The Prompt Isn’t Always the Query: What Conversation Context Means for SEO and GEO (Hamsterdam Research)
By Ethan Lazuk
Last updated:

Welcome to another edition of Hamsterdam Research! 🐹
This is where we look at recent AI research papers to learn what they’re talking about and explore their hypothetical implications for the future of search and SEO/GEO strategies.
This time, we’ll look at a research paper called, “Beyond the Final Prompt: Measuring the Effect of Within-Conversation Context on AI Answers.“
It was published on August 3, 2026, and its author is Benjamin Tannenbaum.
Before we dive in, why should SEO/GEO professionals care about this research paper?
This paper is relevant to SEO/GEO professionals because it challenges a basic assumption behind a lot of AI visibility tracking: that a single prompt can be treated as the query.
The study found that when the same final user message was answered with the preceding conversation (as opposed to without it), the answer changed materially in 44.7% of cases. Those changes included different recommendations, conclusions, clarification behavior, and constraint compliance. For GEO measurement, that matters because brand mentions, citations, and recommendations may depend not just on the prompt being tracked, but on the accumulated context of the conversation leading up to it.
The practical implication for SEO/GEO is that prompt tracking may need to evolve from isolated prompts toward conversational scenarios or request states. A brand might appear for “What’s the best option?” in one conversation and disappear in another conversation because the user previously added a budget, rejected a product, or specified a requirement.
This paper argues that visibility, brand citation, and recommendation measurements can’t always be attributed to the final prompt alone. That doesn’t mean every prompt tracker must replay full conversations, but it does suggest that single-prompt benchmarks give only a partial picture of real AI search behavior.
To start, we’ll review the paper’s abstract.
Here’s the abstract (with my highlights):
“An isolated final user message is often treated as the query in evaluations of AI systems. In a conversation, however, the actionable request may be distributed across preceding turns. We directly test whether that omitted within-conversation context changes answers. For each of 180 English multi-turn conversations sampled from a governed commercial corpus and the public PRISM dataset, we hold the final user message and requested answer model constant while generating three answers: one from the full role-labelled conversation, one from the final message alone, and one from the final message plus a prefix-only reconstruction capped at 160 words. A separately requested judge model evaluates answers under randomized labels. The prespecified primary endpoint is a material difference that could change what the user does, rather than a difference in style or detail. After inverse-probability weighting to the eligible cohorts, the full-conversation and isolated-final answers differ materially in 44.7% of cases (95% bootstrap CI 33.8%–56.1%). Full-conversation answers score 0.49 points higher on a 0–4 request-satisfaction scale (0.32–0.67). Adding the compressed prefix reduces the material-difference rate to 30.8% (20.2%–42.1%), a 13.9-point reduction (4.9%–24.1%), and reduces the mean satisfaction gap to 0.01 points (–0.13). Yet compression is not equivalent to the complete dialogue context: almost one third of answers remain materially different. The primary comparison is stronger in the commercial cohort (68.5%) than in PRISM (35.4%). An order-swapped repeat on 48 cases yields 91.7% agreement and for the primary decision. The results identify the conversation, not its endpoint string, as the defensible query unit for many AI-answer measurements. They concern preceding turns in the same conversation and do not test persistent memory across separate conversations.”
Next, let’s break down the key vocabulary terms from the research paper.
- Within-conversation context: Information contained in earlier turns of the same conversation that helps determine what the user actually means or wants. This is the paper’s central concept.
- Final user message/endpoint string: The last message a user sends. The paper argues this often should not be treated as the complete query because important goals and constraints may have appeared earlier.
- Conversation prefix: All of the conversation turns that come before the final user message. In GEO terms, this is the context that may influence which brands, sources, or recommendations appear in the answer.
- Request state: The accumulated understanding of the user’s goal, constraints, preferences, rejected options, and unresolved questions across a conversation. The paper treats this as more meaningful than looking at the final prompt alone.
- Material difference: A change in the AI answer significant enough to alter what the user might do, such as changing a recommendation, conclusion, constraint compliance, or clarification behavior.
- Request satisfaction: A measure used in the study to judge how well an answer fulfills the user’s actual request. Full-conversation answers scored better on average than answers based only on the final message.
- Compressed reconstruction: A short summary of the important preceding conversation context. It improved results compared with using the final prompt alone, but still produced materially different answers from full conversation context in about 30.8% of cases.
- Conversation as the query unit: Probably the most important GEO phrase in the paper. The author argues that for many AI measurements, the conversation-conditioned request, rather than an isolated prompt, is the more defensible unit of analysis.
Cool. Let’s take a deeper look at the research paper’s contents.
If you want to follow along, you can grab a PDF or the HTML version on arXiv.
The full paper has ten sections.
We’ll summarize each of them below.
1. Introduction
This paper introduces one “deliberately narrow question”: “Holding the final user message constant, how do the preceding turns in the same conversation change the answer?”
In terms of findings:
“The central result is not that every answer needs every earlier token. It is that the endpoint string is often an inadequate experimental unit. In the weighted pooled estimate, final-message isolation materially changes the answer in 44.7% of cases. Compression closes the average satisfaction gap but leaves a 30.8% material-difference rate. A concise state summary can be useful without being interchangeable with the dialogue context that produced it.”
2. Related Work
The paper looks at three domains of related work:
- Conversational information seeking and rewriting
- Context use in language models
- The unit of measurement in AI search
First, conversational information retrieval research has already shown that a user’s information need can develop across multiple turns, rather than exist in a single standalone query. Prior systems such as contextual question rewriting, ConvGQR, and CHIQ try to convert a context-dependent message into a self-contained query that a retriever can use. This paper goes a step further: instead of asking whether context improves the retrieval query, it asks whether including or excluding the earlier conversation actually changes the final AI answer. The paper also points out that summarizing prior context may preserve explicit facts while losing subtler conversational signals, such as whether a suggestion was accepted, rejected, corrected, or left unresolved.
Second, prior work on how language models use conversation history has produced mixed findings. Earlier dialogue systems did not always use previous turns effectively, and newer research suggests models can sometimes perform worse when information is spread across a long conversation. Other work has studied which turns matter most or whether earlier assistant-generated text can actually become distracting. That means more context is not automatically better. The author therefore distinguishes between an answer being different and an answer being better.
Lastly, the paper connects this research to measurement in AI search. Previous research suggests conversational AI can compress multiple search and source-selection actions into one answer, while a user’s short final message can depend on requirements accumulated earlier in the conversation. This paper tests the connection between those ideas. If the same final prompt produces different answers depending on prior context, then metrics such as AI visibility, brand citations, factuality, and recommendations cannot always be attributed to the final prompt alone. The paper also acknowledges that using an LLM to judge the outputs introduces potential bias, so its randomized evaluation setup improves reliability but does not replace human validation.
3. Estimand and Experimental Design
The researcher tested whether conversation history changes an AI answer even when the final prompt stays exactly the same. For each of 180 conversations, he generated three answers: one using the full conversation, one using only the final user message, and one using a short reconstruction of the earlier context. Everything else, including the model and final prompt, was held constant. The compressed version summarized things like goals, constraints, rejected options, corrections, and unresolved questions.
A separate blinded AI judge then compared the answers. A difference only counted as material if it could change the user’s outcome, such as producing a different recommendation, conclusion, constraint handling, clarification, or refusal. (Wording and tone differences didn’t count.) The judge also scored request satisfaction and severity, while the researcher used randomized answer labels, repeat judging, frozen analysis rules, weighting, and bootstrap confidence intervals to improve reliability and reduce bias.
4. Data and Sampling
This section explains where the 180 conversations came from and how the researcher tried to make the sample balanced and appropriate for the experiment. He used two datasets: a private commercial corpus and the public PRISM dataset. After removing duplicates and applying eligibility rules, he had 544 commercial conversations and 1,384 PRISM conversations. To qualify, a conversation needed multiple user turns, at least one assistant response before the final message, a final message between 2 and 96 words, and a total context under about 12,000 tokens. He also excluded sensitive or high-stakes topics.
From those eligible conversations, the researcher selected 90 from each source, for 180 total, and deliberately included a mix of short, medium, and long conversations, as well as cases with different levels of apparent dependence on prior context. Because the two source pools were different sizes, the final statistical results were weighted so the commercial and PRISM data reflected their actual proportions in the eligible population rather than the artificial 50/50 sample. For privacy, the public dataset released from the experiment contains only anonymous IDs, numerical measures, and other text-free features. The actual conversations and model outputs remain private.
5. Implementation
The implementation section explains the technical setup the author used to run the experiment consistently. He used gpt-5.4-mini to generate the reconstructed context and answers, and gpt-5.5 as the judge. Across 180 conversations, he produced 180 context reconstructions, 540 answers, 180 primary judgments, and 48 repeat judgments. Each answer condition was generated once per case, with standardized prompts, schema constraints, retries for technical failures, and protections against conversation text influencing the evaluator.
He also built in reproducibility and quality-control measures, including cryptographic hashes for the input files, configuration, and analysis code, along with unit tests covering label randomization, weighting, bootstrap calculations, lexical distance, judge agreement, and privacy safeguards. In summary, the implementation was highly controlled and auditable, although each condition still relied on only one stochastic model generation per conversation.
6. Results
The main finding is that conversation history often materially changes AI answers. When the model saw the full conversation instead of only the final user message, the answer differed materially in 44.7% of cases. Full-context answers also scored higher for request satisfaction, averaging 3.80 vs. 3.31 on a 0–4 scale. The differences were not just cosmetic: 26.7% involved a changed recommendation or conclusion, 14.4% changed clarification or refusal behavior, and 10.6% changed whether the answer followed the user’s constraints.
A compressed summary of the earlier conversation helped, but it did not fully reproduce the effect of having the complete dialogue. It reduced the material-difference rate from 44.7% to 30.8%, and its average satisfaction score was almost identical to the full-context condition, yet nearly a third of answers still differed in meaningful ways. The effect also varied by dataset: the commercial conversations showed a much larger context effect than PRISM. The blinded judge was fairly consistent across repeated evaluations, with 91.7% agreement on the main comparison, although the paper notes this is not the same as validation by human experts.
7. Interpretation
The interpretation is that the conversation, not just the final prompt, is often the better unit of measurement for AI search. Replaying only the last user message can produce a materially different answer, which matters for GEO because it can change which brand is cited, which product is recommended, or whether the model answers versus asks for clarification. The researcher argues that evaluations should therefore preserve whatever context policy the real system uses, whether that is full history, summarization, retrieval over prior turns, or some other dialogue-state method.
The compressed-context condition shows that good answer quality does not necessarily mean the same answer. A short reconstruction captured enough context to nearly match full-conversation satisfaction scores, but about 30.8% of answers were still materially different. That means GEO benchmarks should distinguish between answer quality and answer identity, and prompt sets based only on final messages should be treated as endpoint-message tests rather than true conversational evaluations. The paper also cautions against assuming one benchmark rate applies universally, since the strength of the context effect varied substantially between the two datasets.
8. Limitations
The main limitations are that the study used only one answer model and one technical setup, so the findings may not generalize to other models, chat applications, system instructions, tool use, or context-management approaches. It also generated only one answer per condition, which means some of the observed differences could reflect normal stochastic variation rather than conversation context alone. In addition, the answers were evaluated by another AI model rather than human judges, so the reliability checks reduce some bias but do not eliminate the possibility of shared model bias or judgment errors.
The study also tested only one 160-word compression method, so it does not show that this is the best way to summarize conversation context. Its datasets and weighting choices also limit how broadly the exact percentages should be applied. Most importantly, the research only examines context within a single conversation. It does not test persistent memory, cross-session personalization, retrieved user data, or other hidden application state. Sensitive and high-stakes topics were excluded as well, so the findings should not be generalized to those areas.
9. Ethics, Privacy, and Reproducibility
This section explains how the researcher handled privacy and reproducibility. He only sent conversation data to the model after receiving explicit authorization, and he did not publish raw conversations, model outputs, or identifying source information. The public artifacts contain only hashed case IDs, broad source labels, sampling information, and numerical results, while the commercial dataset remains private and PRISM stays subject to its original usage rules.
At the same time, he released enough of the experimental setup to make the study auditable. This allows others to inspect how the sampling and analysis were performed without exposing private conversations, although fully reproducing the model outputs would still require authorized access to the original data and model interfaces.
10. Conclusion
Let’s examine this in full:
“The final user message is not generally the whole query. With that message held constant, removing preceding turns materially changes 44.7% of answers in the weighted eligible cohorts and lowers judged request satisfaction by nearly half a point on a 0–4 scale. A compressed reconstruction recovers most of the average satisfaction gap and reduces the material-change rate by 13.9 points, but it remains behaviorally different from the full conversation in 30.8% of cases.
The appropriate lesson is methodological. Evaluations of conversational AI must specify the conversation prefix, dialogue state, or accumulated request context that conditions an answer. Replaying endpoint strings alone measures a different task. Persistent memory across separate conversations remains a separate research question.”
So, why should SEO/GEO professionals care about “Beyond the Final Prompt: Measuring the Effect of Within-Conversation Context on AI Answers”?
Here are five takeaways:
1. A prompt is not always the full query.
In conversational AI, the user’s real information need can be built across several turns, so evaluating only the final message can miss important context.
2. Conversation context can materially change AI visibility.
In the study, the same final message produced a materially different answer 44.7% of the time when earlier conversation history was included.
3. Recommendations, conclusions, and constraint handling can change.
This matters for GEO because a different conversational path could lead to a different brand recommendation, citation, or product selection.
4. Single-prompt tracking is useful, but incomplete.
Prompt-tracking tools may measure endpoint-message visibility rather than the full range of outcomes users could see in realistic multi-turn conversations.
5. Summarized context helps, but does not fully solve the problem.
A compressed version of the prior conversation improved answer quality, yet 30.8% of responses were still materially different from those generated with full context.
Outro
I hope you’ve enjoyed this edition of Hamsterdam Research! 🐹
Feel free to comment below or contact me with your feedback.
Stay tuned for another new article, hopefully next week, or check out related research posts below.
Until next time, enjoy the vibes:
Thanks for reading. Happy optimizing!
Related research articles:
Exploring “Scalable In-context Ranking with Generative Models,” a Google Research Paper, & Why SEO/GEO Professionals Should Care (A Hamsterdam Research Post)
Exploring “Scalable In-context Ranking with Generative Models,” a Google Research Paper, & Why SEO/GEO Professionals Should Care (A Hamsterdam Research Post) By Ethan Lazuk Last updated: Welcome to a new edition of Hamsterdam Research! 🐹 If you’re new here, this is where we look at recent AI research papers to learn just what the heck…
Summer “SLaM” & “CoSMo” Kramer: Investigating “Compressing Search with Language Models,” a Google Research Paper, & Why SEOs Should Care (Probably)
We’ll explore SLaM and CoSMo from a Google paper, “Compressing Search with Language Models,” and implications for SEOs in this Hamsterdam Research post.
Exploring “GuidedRAG: Semantic Steering of Retrieval-Augmented Generation” and Why SEO/GEO Professionals Should Care (Hamsterdam Research)
Exploring “GuidedRAG: Semantic Steering of Retrieval-Augmented Generation” and Why SEO/GEO Professionals Should Care (Hamsterdam Research) By Ethan Lazuk Last updated: Welcome to a new edition of Hamsterdam Research! 🐹 This is where we look at recent AI research papers to learn just what the heck they’re talking about and explore their hypothetical implications for the…
Leave a Reply