Summarizing “System Attribution in LLM Brand Recommendations: Single Responses Identify the System, Aggregated Brand Profiles Do Not Transfer”: A Hamsterdam Research Post
By Ethan Lazuk
Last updated:

Welcome to another edition of Hamsterdam Research! š¹
This is where we look at recent AI research papers to learn what they’re talking about and explore their hypothetical implications for the future of search and SEO/GEO strategies.
This time, we’ll look at research called, “System Attribution in LLM Brand Recommendations: Single Responses Identify the System, Aggregated Brand Profiles Do Not Transfer.”
It was submitted to arXiv on September 23, 2026, and its author is Dmitrij Żatuchin.
First off, why should SEO/GEO professionals care about this research?
This paper asks a question underlying GEO visibility reporting, which is that when we say a model “prefers,” mentions, or recommends certain brands more often, are we actually measuring a stable characteristic of the AI system or are we partly measuring the prompts, query category, and testing setup we used?
Much of AI visibility tracking today works by aggregating responses into metrics like brand mentions, share of voice, sentiment, citation frequency, or recommendation rates. In this study, those aggregated brand-behavior profiles didn’t transfer well between query domains. Meaning, a model could appear to behave one way for gift recommendations (like “What should I get my boyfriend for Valentineās Day?”) and quite differently for category-ownership queries (like “Which companies are most associated with fintech?”).
One interesting contrast found was that individual responses themselves were highly identifiable by system (things like formatting, layout, punctuation, etc.), but the higher-level brand behavior that GEO practitioners normally care about was much less stable. Thus the authors found the model may have a recognizable “way of answering” without having an equally stable “way of recommending brands.”
The main takeaway isn’t “AI visibility scores are useless” but rather that AI visibility is conditional. When we say a score for “Gemini visibility” we really may mean “Gemini visibility for this particular collection of prompts, in these categories, through this testing setup, at this point in time.”
Let’s start by reviewing the paper’s abstract.
This is worth a read in full, particularly the first half:
“Audits of AI visibility summarise the brand recommendations of deployed language models into per-system profiles. This paper asks whether such a profile describes the system, and puts the question to one corpus at two units of observation. The corpus holds 6,475 stored responses, 6,324 of them analysable, collected between December 2025 and February 2026 from five deployed endpoints across gift-recommendation, corporate-reputation and category-ownership queries. The collection harness cut many answers short: 83.1% of the Gemini 3 Flash answers in category ownership end mid-sentence under a 1,024-token output cap. With every answer cut to its first 800 characters, so that where an answer stops carries no information, a character -gram classifier cross-validated by prompt attributes a single response to GPT-5.2, Gemini 3 Flash, Gemini 3 Flash with search, Grok or Perplexity sonar-pro with 97.84% accuracy (5,028 responses, 383 prompts, majority class 31.5%, 30 split seeds); full-length text gives 98.14%. Length alone then falls to the majority rate, 24 formatting statistics reach 95.79%, and masking brand names and capitalised tokens leaves 97.72%. Held-out query conditions inside the same harness keep 97.43% weighted by size and 88.0% unweighted; in the one held-out condition that changes the harness, a retrieval-grounded arm, no Grok answer is attributed to Grok (0/120). Aggregated into 50 model-by-domain-by-condition units, twelve behavioural features separate four systems at 66.53% under grouped cross-validation, against a label-permutation null with mean 33.71% and 95th percentile 46.0%; two of the four systems occur in one domain only, and the feature carrying the separation tracks domain and repetition count. Across domains the aggregate profile fails: a forest trained on ten category-ownership units assigns all 22 gift units to the wrong system, consistent with a reversal in brand volume between the two systems (8.41 against 0.94 brands per response in gifts, 3.01 against 3.91 in category ownership), while single responses transfer at 89.92% balanced accuracy. The surface form of an answer carries the system across the query domains tested; aggregated brand behaviour does not, and the uncrossed design cannot separate the system from the domain or the harness.”
Next, let’s break down the paper’s key vocabulary terms:
- Deployed endpoint: The specific AI system or configuration producing a response, such as GPT-5.2, Gemini 3 Flash, Gemini with search, Grok, or Perplexity sonar-pro.
- Collection harness: The technical setup used to collect responses, including the scripts, instructions, search settings, completion limits, and other conditions surrounding the model.
- System attribution: Determining which AI system produced a particular response based on characteristics of that response.
- Character n-gram: A short sequence of characters used as a text feature. The researchers used character n-grams of two to four characters to identify patterns in things like punctuation, capitalization, list formatting, and wording.
- Aggregate unit: A collection of many responses grouped by model, query domain, and condition. This is closer to the level at which an AI visibility platform might calculate an overall visibility score or model profile.
- Brand abstinence rate: The proportion of responses in which the AI does not mention any detected brand.
- Brand concentration: The extent to which an AI system’s brand mentions are concentrated among a small number of brands rather than distributed across many. The researchers measured this using metrics including the Gini coefficient, Shannon entropy, and top-brand share.
- Response consistency: How similar the model’s answers are when the same prompt is run repeatedly, measured using the similarity between repeated responses.
- Cross-domain transfer: Testing whether a pattern learned from one type of query, such as gift recommendations, still works on a different type, such as category-ownership queries.
- Calibration: How closely a model’s stated confidence corresponds to how often its prediction is actually correct. The paper measures this using Expected Calibration Error (ECE).
Esta bien. Letās take a deeper look at the research paper’s contents now.
If you want to follow along, you can grab a PDF or the HTML version on arXiv.
The paper has 7 main sections.
Weāll focus on the most relevant ones below.
Introduction
“Generative engine optimisation treats those answers as a surface that content can be shaped for and measured on,” the author writes. “This paper starts from that premise and tests the assumption beneath any per-system measurement of it: that an answer can be attributed to the system that produced it, and that what the measurement observes is a property of that system. The query domain and the collection harness are the two competing sources this corpus lets us examine.”
“The test uses one corpus at two units of observation, and the two units give different answers,” he writes.
“The first unit is the single stored response.”
“The second unit is the model-by-domain-by-condition aggregate, the unit an AI-visibility report summarises.”
Results
The biggest finding is that individual AI responses have surprisingly strong āfingerprints,ā but aggregated brand-behavior profiles are much less stable.
- The researcher could usually tell which AI system wrote a response. Across five systems, he identified the source model correctly about 98% of the time. Even when every answer was shortened to the first 800 characters, accuracy stayed at 97.84%.
- It was mostly the way the models wrote, not the brands they mentioned, that gave them away. Formatting characteristics alone reached 95.79% accuracy. Things like sentence length, citations, numbers, lists, punctuation, and layout carried a lot of the identifying signal. Removing brand names and capitalized words barely changed the results. Simply looking at answer length, meanwhile, was basically useless once responses were standardized.
- That fingerprint generally held up when the researchers changed the prompts but kept the same testing setup. Across 16 held-out conditions, identification accuracy was 97.43% when weighted by the number of responses. But performance wasn’t equally strong everywhere: some Valentineās Day gift categories fell to around 60%ā77%.
- The aggregated brand profiles were much less dependable. When the researchers summarized many responses into features such as brand counts, concentration, sentiment, consistency, and citation behavior, they could identify the system only about 66.5% of the time. That’s better than the study’s appropriate baseline, but dramatically weaker than the ~98% result for individual responses.
- The aggregated profiles basically fell apart when moved to another query domain. This is probably the most important GEO result.
Discussion
“A single response identifies its endpoint at 97.84% over five systems with truncation removed,” meanwhile “An aggregate of tens to hundreds of the same responses, summarised into brand volume, concentration, sentiment and source behaviour, separates four systems at 66.53% inside the window and inverts across domains,” the author writes.
He continues: “The brand-volume ordering of two systems flips between gift prompts and category-ownership prompts, and a profile built on that ordering flips with it.”
And finally: “A measurement that reports per-system brand behaviour reports at the aggregate unit. In this corpus a brand-behaviour profile measured in one query domain describes that domain, and carrying it to another domain needs evidence.”
Conclusion
“From a corpus of 6,324 analysable responses produced by five deployed endpoints in three query domains, a character -gram classifier grouped by prompt attributes a single answer to its endpoint at 97.84% on 5,028 responses cut to a common 800-character prefix (98.14% on the stored text), keeps 97.43% on held-out query conditions within the same harness (88.0% unweighted, the small Valentineās conditions at 60.0 to 76.67%), and is close to calibrated there. Twenty-four formatting statistics reach 95.79% on the same prefixes. In the one held-out condition that changes the harness, Perplexity is attributed correctly and Grok is not.
The same corpus aggregated into 50 model-by-domain-by-condition units and summarised by twelve behavioural features separates four systems at 66.53%, against a structure-preserving permutation null with 95th percentile 46.0%, and one feature that tracks domain and repetition count carries most of that. Moved between the gift and category-ownership domains, the aggregate classifier gets 4 of 32 units right, while single responses keep 89.92% balanced accuracy on the same two-system question.
The shape of an endpointās answers carries its identity across the query domains tested. Its aggregated brand behaviour does not, and this design cannot say how much of that behaviour belongs to the system and how much to the domain or the harness. An AI-visibility measurement built on aggregated brand behaviour is bounded accordingly.”
Knowing what we do now, why should SEO/GEO professionals care about “System Attribution in LLM Brand Recommendations: Single Responses Identify the System, Aggregated Brand Profiles Do Not Transfer”?
This paper challenges a common assumption in AI visibility measurement that the brand behavior we observe from a model is a stable property of that model.
The study found that individual responses from systems like GPT-5.2, Gemini, Grok, and Perplexity were highly identifiable. In other words, the systems had recognizable response patterns. But the aggregated brand behaviors GEO practitioners often track, such as how many brands a model mentions, how concentrated those mentions are, how consistent the recommendations are, and how often citations appear, were much less stable.
The most important result is that those brand-behavior profiles did not transfer well across different kinds of queries. Gemini mentioned far more brands than GPT-5.2 in the gift-recommendation domain, but that relationship reversed for category-ownership queries.
GEO platforms often roll many prompts into a single visibility score or make statements such as āGemini favors more brandsā or āChatGPT has higher brand concentration.ā This paper suggests those conclusions may describe the prompt set and query domain as much as the model itself.
There is also a measurement-methodology lesson here. The researcher found that things like formatting, sentence structure, citations, response length, completion limits, and the collection setup itself can influence what appears to be āmodel behavior.ā In one test where the harness changed, Perplexity remained identifiable while Grok did not, showing that even the technical setup around the model can affect the results.
Outro
For SEO/GEO professionals, the practical takeaway is that AI visibility should be treated as conditional, not universal. A visibility result is better understood as “this brandās visibility in this model, for this type of query, using this prompt set, under this measurement setup, at this point in time.”
I hope you’ve enjoyed this edition of Hamsterdam Research! š¹
Feel free to comment below or contact me with your feedback.
Stay tuned for another new article, hopefully next week, or check out related research posts below.
Until next time, enjoy the vibes:
Thanks for reading. Happy optimizing!
Related research articles:
Evaluating “The Fellowship of the Query: Learning Retrieval Actions” for SEO/GEO Insights: A Hamsterdam Research Post
Evaluating “The Fellowship of the Query: Learning Retrieval Actions” for SEO/GEO Insights: A Hamsterdam Research Post By Ethan Lazuk Last updated: Welcome to another edition of Hamsterdam Research! š¹ This is where we look at recent AI research papers to learn what they’re talking about and explore their hypothetical implications for the future of searchā¦
Summarizing “Query Implied Generative Engine Optimization”: A Hamsterdam Research Post
Summarizing “Query Implied Generative Engine Optimization”: A Hamsterdam Research Post By Ethan Lazuk Last updated: Welcome to another edition of Hamsterdam Research! š¹ This is where we look at recent AI research papers to learn what they’re talking about and explore their hypothetical implications for the future of search and SEO/GEO strategies. This time, we’llā¦
Summarizing “Conversational Capture: A Trajectory-Level Framework for Evaluating Generative Engine Optimization in Multi-turn Human-Agent Interaction”: A Hamsterdam Research Post
Summarizing “Conversational Capture: A Trajectory-Level Framework for Evaluating Generative Engine Optimization in Multi-turn Human-Agent Interaction”: A Hamsterdam Research Post By Ethan Lazuk Last updated: Welcome to another edition of Hamsterdam Research! š¹ This is where we look at recent AI research papers to learn what they’re talking about and explore their hypothetical implications for theā¦
Leave a Reply