Reviewing “Evaluating Brand Retrieval and Ranking in Large Language Model Recommendations” for SEO/GEO Insights: A Hamsterdam Research Post
By Ethan Lazuk
Last updated:

Welcome to another edition of Hamsterdam Research! š¹
This is where we look at recent AI research papers to learn what they’re talking about and explore their hypothetical implications for the future of search and SEO/GEO strategies.
This time, we’ll look at research called, “Evaluating Brand Retrieval and Ranking in Large Language Model Recommendations.”
It was submitted to arXiv on September 14, 2026, and its authors are Edward Malthouse, Kun-Yu Lee, Jing Yang, Sanchary Pal, and Xueyan Feng.
First off, why should SEO/GEO professionals care about this research?
This paper attempts to answer a practical question for SEO/GEO professionals, which is what actually determines whether a brand gets recommended by an LLM, and how should we measure that visibility.
The researchers treat LLM recommendations as stochastic retrieval-and-ranking systems, where the same prompt can return different brands in different orders. This means checking a single ChatGPT or Gemini response is not a reliable way to measure SEO/GEO visibility. Instead, the paper proposes repeated sampling and introduces two metrics:
- Brand Recommendation Probability (BRP@k) measures whether a brand appears in the LLMās recommendation set (āHow often am I recommended?ā).
- Mean Reciprocal Rank (MRR@k) measures how prominently the brand appears (āHow high am I recommended when I appear?ā).
The researchers’ findings challenge some easy assumptions about AI visibility. For example, large, established brands were sometimes completely omitted from recommendations, and traditional brand salience did not consistently predict which brands surfaced. Instead, recommendation prominence was associated more strongly with broader marketplace visibility, particularly Google search interest and online brand conversation, although the authors were careful to say these are correlations, not proof that increasing search demand will directly cause better LLM rankings.
The paper is also relevant to SEO/GEO because it shows that query context changes brand retrieval. Brands that were invisible for generic category prompts could begin appearing when users described specific needs, constraints, use cases, and preferences. And when researchers supplied cues closely matching a brand’s positioning, recommendation rates increased dramatically.
In other words, SEO/GEO strategies may be about more than simply getting a model to “know” a brand. Visibility may also depend on whether the model has strong enough associations between that brand and the situations in which someone should consider it, i.e., a relevant consumer need is expressed. That is a more sophisticated way of thinking about GEO than simply asking, āDoes ChatGPT mention my brand?ā
Let’s start by reviewing the paper’s abstract.
Here’s the abstract, which is worth a read in full:
“Large language models (LLMs) are increasingly used for product recommendation, but evaluating their recommendations presents challenges that differ from conventional information retrieval and recommender systems. LLMs can generate recommendations without an explicit candidate set, and repeated responses to the same query can produce different brands and rankings. We introduce a framework for evaluating open-ended LLM brand recommendations that defines the competitive set independently of model outputs and estimates recommendation prevalence and prominence through repeated sampling. We operationalize these constructs using Brand Recommendation Probability (BRP@) and Mean Reciprocal Rank (MRR@), and apply the framework to six LLMs across five product categories. Category-only queries reveal substantial omission of established brands and limited evidence that recommendation prominence follows conventional brand popularity. Instead, prominence is associated with broader marketplace-visibility signals, particularly search interest and online brand conversation. Needs-based queries show that contextualizing usersā goals and constraints changes which brands are retrieved, while diagnostic positioning probes demonstrate that brands omitted from ordinary recommendations can remain conditionally retrievable when distinctive cues are supplied. These findings highlight the need to evaluate LLM recommendation as a stochastic retrieval-and-ranking process rather than from individual generated lists. We provide open-source software and data to support reproducible evaluation of LLM-generated brand recommendations.”
Next, let’s break down the paper’s key vocabulary terms:
- Generated Set: The small set of brand alternatives an LLM constructs in response to a userās prompt. Unlike a traditional search-results page, the set can change across repeated queries.
- Brand Recommendation Probability (BRP@k): The probability that a brand appears within the first k recommendations across repeated LLM responses. It measures recommendation prevalence.
- Mean Reciprocal Rank (MRR@k): A metric for recommendation prominence that gives more weight to brands appearing higher in the recommendation list and assigns zero when a brand is absent.
- Recommendation Prevalence: How frequently a particular brand enters the LLM-generated recommendation set across repeated prompts.
- Recommendation Prominence: How visible or highly positioned a brand is within LLM recommendations, rather than simply whether it appears at all.
- Conditional Retrievability: Whether an LLM can retrieve a brand when it is given distinctive cues associated with that brandās positioning, even if the brand does not normally appear in recommendations.
- Needs-Based Prompt: A prompt that includes a consumerās goals, constraints, preferences, budget, intended use, or other context rather than merely naming a product category.
- Positioning Probe: A diagnostic prompt containing distinctive cues related to a brandās intended positioning, used to test whether the LLM can retrieve and associate the brand with those characteristics.
- Marketplace Visibility: The broader information environment surrounding a brand, represented in the study by signals such as advertising, search interest, news coverage, online conversation, and Wikipedia activity.
- Choice Architecture: The way the presentation, selection, and ordering of alternatives can shape consumer decisions. The authors describe LLMs as emergent choice architects because they construct the alternatives shown to users at the moment of recommendation.
Good stuff. Letās take a deeper look at the research paper’s contents now.
If you want to follow along, you can grab a PDF or the HTML version on arXiv.
The paper has five main sections.
Weāll focus on the most relevant ones below.
Introduction
By giving product recommendations, an LLM “becomes an intermediary between consumer needs and brands, potentially influencing which brands enter consideration and the order in which consumers encounter them,” write the authors. “A brand omitted from an LLMās recommendations may never be considered, even when it fits the consumerās needs.”
They next talk about marketers’ need for measurement:
“This new form of gatekeeping creates a measurement problem for marketers. The set of brands an LLM may recommend is neither fixed nor directly observable, and repeated requests can produce different recommendations and rankings. Brand managers therefore need methods for determining whether their brands are recommended, how frequently and prominently they appear, which competitors appear instead, and how recommendations change when consumers provide information about their needs.”
They identify two important issues, “First, LLM recommendations may favor a relatively small subset of brands.”
“Second,” they write, “consumers can communicate their goals, preferences, and constraints directly to LLMs [and] Such needs-based prompts provide information that may help LLMs match consumer needs with brands, potentially bringing otherwise omitted brands into the recommendation set.” They continue: “This possibility is especially important for firms that invest in differentiated positions intended to associate their brands with particular consumer needs, use cases, and value propositions.”
Next they describe their measurements and methodology:
“We first define the competitive set independently of the LLM, making it possible to identify not only recommended brands but also omitted ones. We then define Brand Recommendation Probability (BRP) to measure recommendation prevalence and adapt Mean Reciprocal Rank (MRR) to measure recommendation prominence. Using repeated recommendations from multiple LLMs across five product categories, we examine category-only recommendations and explore marketplace factors associated with recommendation prominence. We then use needs-based prompts to examine whether brands surface when consumers express relevant goals and constraints, followed by diagnostic positioning probes to assess whether omitted brands can be retrieved when supplied with distinctive cues consistent with their intended positioning.”
Conceptual Background and Framework
“Choice architecture describes how the environment in which alternatives are presented, including which options appear, how many appear, and how they are ordered, shapes choice,” they write.
“Retailers determine which brands appear on physical shelves and where they are placed; search engines rank pages and sell sponsored positions; and recommender systems select and rank alternatives from platform inventories. Each environment has identifiable mechanisms through which brands gain exposure and established metrics for monitoring that exposure.”
They next talk about how LLMs differ as choice architecture:
“LLMs differ from these environments in an important respect: rather than selecting alternatives from a fixed catalog or inventory, they can generate a set of brand recommendations in response to a consumer’s prompt. We therefore describe LLMs as emergent choice architects: they construct a small set of alternatives at the moment of recommendation, even though no category manager or other actor has explicitly selected the brands that will appear. For marketers, this makes brand exposure less transparent. There is no shelf allocation, search ranking, or platform inventory to inspect; instead, the brands that surface must be observed in LLM responses, with the model itself acting as gatekeeper to that visibility.”
They next talk about generated sets versus consideration sets:
“Consumers can describe a need, goal, or problem in natural language and ask the LLM what they should buy. The model then returns a small set of brands on the consumerās behalf. We call this a generated set: a compact set of alternatives constructed by the LLM in response to the consumerās prompt. Unlike an assortment or catalog-based top- list, membership and ordering in a generated set may vary across repeated queries.”
They continue: “This perspective creates two fundamental measurement questions for marketers: How frequently does a brand enter the generated set, and how prominently does it appear when recommended?”
Next they discuss their metrics:
“We adapt metrics from information retrieval and recommender systems to capture these dimensions. Brand Recommendation Probability (BRP@) measuresĀ prevalence, or the probability that a brand appears among up to recommendations. Mean Reciprocal Rank (MRR@) measures prominence by considering both whether the brand appears and its position in the recommendation list. Because LLM responses can vary across repeated requests, both quantities are estimated from replicated queries. Formal definitions and estimation procedures are provided in the Framework subsection below.”
They then discuss their framework: “Building on the generated-set perspective, we propose a five-step framework for evaluating LLM brand recommendations.”
- “Define. The first step is to establish the competitive set independently of the LLM responses.”
- “Measure. The second step usesĀ category-only promptsĀ to establish baseline brand recommendations, quantifying recommendation prevalence with Brand Recommendation Probability (BRP@) and recommendation prominence with Mean Reciprocal Rank (MRR@).”
- “Explore. Once differences in recommendation prominence have been established, external marketplace data can be used to explore characteristics associated with those differences.”
- “Match. Category-level prominence alone does not establish whether an LLM recommends appropriate brands for consumers of a particular focal brand. The sampling procedure used in the Measure stage can therefore be repeated using needs-based prompts that articulate goals, constraints, preferences, and intended uses for a target consumer of the focal brand.”
- “Diagnose. When a brand fails to surface for consumer needs it is intended to serve, diagnostic probes can be used to investigate the LLM’s representation of the brand and its positioning.”
Results
“Three findings stand out,” they write:
“First, many large, established brands receive no recommendations at all. …
Second, the figures provide only limited evidence of the popularity bias documented in the marketing and recommender-systems literature. …
Third, BRP@5 and MRR generally tell a consistent story, but MRR provides additional discrimination when brands have similar recommendation prevalence.”
“If popularity does not explain recommendations,” they write, “a key question is: what does?”
“We use the five marketplace-visibility measures described previously (advertising expenditures, press mentions, online brand conversation, search interest, and Wikipedia page views) to explore their associations with recommendation prominence. …
First, there are substantial correlations between many of the explanatory variables, indicating multicollinearity. … the dominant first dimension suggests that the five measures share a substantial common component that we interpret as marketplace visibility. …
Moreover, the causal ordering among the marketplace measures themselves is unclear: online conversation may stimulate search, search may stimulate conversation, or both may respond to common events. …
They next talk about needs-based prompts: “Rather than asking only for brands in a product category, needs-based prompts describe a plausible consumer, use case, goals, and constraints, allowing the LLM to match those needs against its knowledge of competing brands and their positioning.”
“These results raise an important question: are Craftsman and L.L.Bean omitted because the LLMs lack knowledge of their positioning, or because relevant brand knowledge is not activated by ordinary needs-based prompts? To investigate these possibilities, we created moreĀ diagnostic positioning probesĀ that incorporated language drawn directly from the brandsā own marketing materials. We call them probes rather than prompts because they are intentionally unrealistic as consumer requests: a typical shopper seeking advice about a cordless drill or hiking jacket would be unlikely to reproduce detailed brand-positioning language without first researching the market. When supplied with these diagnostic cues, however, BRP@5 increased to 81.3% for Craftsman and 88.5% for L.L.Bean. These results demonstrate that the brands are retrievable when the LLMs are given sufficiently diagnostic cues consistent with their intended positioning. They therefore suggest that low recommendation rates under category-only and needs-based prompts cannot be attributed simply to an absence of brand knowledge. At the same time, conditional retrievability does not establish that the LLMs hold complete or accurate representations of either brandās positioning. Additional probes could systematically test whether the attributes, points of difference, target users, and use cases that LLMs associate with a brand correspond to its intended positioning.”
Discussion
The first part of their discussion section is worth quoting in full:
“Our empirical analyses reveal several patterns in LLM brand recommendations. First, category-only prompts across six LLMs frequently omit large, established brands entirely, indicating that conventional marketplace presence does not guarantee inclusion in LLM-generated recommendation sets. Second, we find very limited evidence of the popularity bias documented for conventional recommender systems: salient brands are sometimes recommended prominently, but the relationship varies substantially across categories. In cordless drills and hiking jackets, recommendations instead tend to favor higher-end or premium brands while neglecting more mass-market alternatives, although this pattern does not generalize across categories. Third, recommendation prominence is systematically associated with observable marketplace signals. Google search interest provides the strongest and most robust predictive signal, followed by online brand conversation, while news mentions, advertising expenditures and Wikipedia page views contribute relatively little once the correlated measures are considered jointly. These relationships are exploratory and should not be interpreted causally. Finally, our needs-based examples demonstrate that providing information about consumersā goals and constraints can cause omitted brands to enter the recommendation set, while more diagnostic positioning probes show that low recommendation rates do not necessarily reflect a lack of LLM knowledge about the brand. Taken together, the findings suggest that LLM brand recommendation reflects both the visibility of brands in the marketplace and the extent to which available prompt information activates brand associations relevant to the consumerās needs.”
Knowing what we do now, why should SEO/GEO professionals care about “Evaluating Brand Retrieval and Ranking in Large Language Model Recommendations”?
Instead of checking a single response, the paper argues that AI recommendations should be measured across repeated prompts using metrics like Brand Recommendation Probability (BRP) and Mean Reciprocal Rank (MRR) to capture how often a brand appears and how prominently it is ranked.
The findings also show that being a well-known brand does not guarantee AI visibility. Search interest, online brand conversation, user needs, and brand positioning were more closely associated with whether and how brands surfaced. For SEO/GEO, that means the key question is not just āDoes the model know my brand?ā but āWhen and how reliably does the model retrieve and recommend it?ā
Outro
I hope you’ve enjoyed this edition of Hamsterdam Research! š¹
Feel free to comment below or contact me with your feedback.
Stay tuned for another new article, hopefully next week, or check out related research posts below.
Until next time, enjoy the vibes:
Thanks for reading. Happy optimizing!
Related research articles:
How Many Times Should You Ask an AI Assistant the Same Question? Research Into Measuring AI Visibility (Hamsterdam Research)
How Many Times Should You Ask an AI Assistant the Same Question? Research Into Measuring AI Visibility (Hamsterdam Research) By Ethan Lazuk Last updated: Welcome to another edition of Hamsterdam Research! š¹ This is where we look at recent AI research papers to learn what they’re talking about and explore their hypothetical implications for theā¦
Does Language Change What AI Search Considers Authoritative? Lessons From 1,920 AI Overview Queries (Hamsterdam Research)
Does Language Change What AI Search Considers Authoritative? Lessons From 1,920 AI Overview Queries (Hamsterdam Research) By Ethan Lazuk Last updated: Welcome to another edition of Hamsterdam Research! š¹ This is where we look at recent AI research papers to learn what they’re talking about and explore their hypothetical implications for the future of searchā¦
What Does Your GEO Visibility Score Actually Measure? Why Prompt Selection Matters (Hamsterdam Research)
What Does Your GEO Visibility Score Actually Measure? Why Prompt Selection Matters (Hamsterdam Research) By Ethan Lazuk Last updated: Welcome to another edition of Hamsterdam Research! š¹ This is where we look at recent AI research papers to learn what they’re talking about and explore their hypothetical implications for the future of search and SEO/GEOā¦
Leave a Reply