How Many Times Should You Ask an AI Assistant the Same Question? Research Into Measuring AI Visibility (Hamsterdam Research)

By Ethan Lazuk

Last updated:

A representation of the dice roll method for AI visibility measurement.

Welcome to another edition of Hamsterdam Research! 🐹

This is where we look at recent AI research papers to learn what they’re talking about and explore their hypothetical implications for the future of search and SEO/GEO strategies.

This time, we’ll look at research called, “The Dice Roll Method: A Standardized Protocol for Repeated-Query Auditing of Large Language Model Brand Recommendations.”

It was submitted to arXiv on September 3, 2026, and its author is Dmitrij Żatuchin.

In short, why should SEO/GEO professionals care about this research?

AI visibility is inherently variable — the same prompt can produce different brand recommendations across repeated runs, meaning a single query (or even a small handful of queries) may give a misleading picture of how visible a brand is in actuality.

The author’s Dice Roll Method proposes a more rigorous way to measure that AI visibility variability by repeating identical prompts, evaluating stability across multiple metrics, and determining how many iterations are needed before the results become reasonably reliable.

The takeaway is not to “run every prompt 15 times.” While the authors identify exploratory, confirmatory, and rigorous iteration ranges, their validation shows that fixed iteration counts do not transfer perfectly across datasets. The broader lesson is to treat AI visibility as a measurement problem — test variability first, then determine how many repetitions are needed for the level of confidence you want.

To elaborate further, AI visibility should probably be thought of as a distribution, not a ranking.

A traditional search result can often be inspected at a particular moment and described as “Brand X ranks at position #3, but with stochastic LLM recommendations, one run may mention Brand X, another may omit it, another may recommend different competitors, and another run may frame the category differently.

The author describes repeated querying as analogous to repeatedly rolling a die in order to estimate the underlying distribution rather than treating any single roll as representative.

If traditional SEO measurement asks “Where did I rank?” then AI visibility measurement may increasingly need to ask “Across repeated generations, how often, how consistently, and under what conditions did I appear?”

Let’s start by reviewing the paper’s abstract.

Here’s the abstract (with my highlights):

“Background: Researchers increasingly use repeated identical prompts to audit stochastic variation in large language model (LLM) brand recommendations, yet no standardized protocol exists for setting iteration counts, selecting stability metrics, or establishing reliability thresholds under the conditional, non-Gaussian structure of autoregressive text generation. Objective: We formalize the Dice Roll Method as a reusable protocol for repeated-query auditing of LLM brand recommendations, grounded in an explicit generative model of temperature-scaled nucleus sampling. Methods: Total response variance is decomposed into token-level sampling, prompt-phrasing, run-to-run, and model-version components. The analytical stack uses a negative-binomial generalized linear mixed model that treats iterations as repeated measures nested within prompt, model, and language; Cliff’s δ as the primary distribution-free effect size; moving-block and parametric bootstrap that preserve the conditional dependence of sequential observations; simulation-based power via simr; a generalizability-theory reliability decomposition; and Kolmogorov–Smirnov and Population Stability Index drift diagnostics on pinned model snapshots. We reanalyse five brand-recommendation auditing studies (approximately 190,000 observations, three to five LLMs, 270+ brands, 6 languages, iteration counts from 5 to 40). Results: Three tiers of iteration guidance emerge from the generalizability-theory D-study: exploratory (n=5, G=0.58), confirmatory (n=10, G=0.74), and rigorous (n=15, G=0.81), with targets tied to Cliff’s δ thresholds and generalizability coefficients. Count-based, set-based, embedding-based, and fairness-adjusted (PASOR) metric families are complementary in a descriptive reading of the bootstrap-corrected Spearman correlation matrix, motivating a compact metric battery over single indicators. A pre-registered external validation on three independent corpora (Motoki et al.’s 100-round political-bias data, Rozado’s 24-model test sweep, and the llm-stability benchmark) reproduces the D-study reliability prediction in 37 of 39 cells with no failures and the n=5 power value to two decimals; the fixed iteration tiers do not transfer, supporting a pilot-then-solve reading of the guidance. Conclusion: The protocol gives repeated-query auditing of LLM brand recommendations a statistically principled footing that holds under the conditional dependencies and non-Gaussian distributions that characterize real autoregressive generation.“

Next, let’s break down the paper’s key vocabulary terms:

  • Repeated-Query Auditing: Repeating the same prompt multiple times and aggregating the outputs to measure how consistently an LLM recommends brands. This is the core behavior the Dice Roll Method is designed to standardize.
  • Stochasticity / Non-Determinism: The fact that an LLM can produce different outputs from the same prompt because generation is probabilistic rather than perfectly deterministic.
  • Iteration Count: The number of times the same prompt is repeated in an audit. The paper focuses heavily on how many iterations are needed before a measurement becomes sufficiently reliable for different goals.
  • Generalizability Theory (G-Theory): A reliability framework the authors use to estimate how much of the observed variation comes from the thing being measured versus noise from iterations, prompts, and models.
  • Generalizability Coefficient (G): A reliability score derived from G-Theory. Higher values indicate that repeated measurements are more dependable; the paper uses it to help define exploratory, confirmatory, and rigorous iteration ranges.
  • Cliff’s Delta (δ): A distribution-free effect-size measure the paper uses instead of relying on normal-distribution assumptions. It estimates how strongly two groups differ without assuming the data are Gaussian.
  • PASOR (Prompt-Adjusted Share of Recommendation): A visibility metric that measures a brand’s share of recommendation while adjusting across prompts, intended to provide a fairer view of brand visibility.

Awesome. Let’s take a deeper look at the research paper’s contents now.

If you want to follow along, you can grab a PDF or the HTML version on arXiv.

The paper has eight main sections.

We’ll focus on the ones most relevant to AI visibility measurement below.

Introduction

The introduction is best laid out using excerpts from the paper:

“The stochastic nature of large language model (LLM) outputs presents a fundamental measurement challenge for repeated-query auditing of LLM brand recommendations. Even at low temperature settings (t=0.3), identical prompts produce varying responses across repetitions [1, 2]. … For researchers seeking to make claims about systematic patterns in LLM brand recommendations, whether measuring gender disparities in recommended brands, evaluating corporate reputation sourcing, or mapping competitive category ownership, the question of how many repetitions are needed to produce reliable measurements is both urgent and largely unanswered.”

“The technique of querying an LLM with the same prompt multiple times and aggregating the results has been adopted independently by several research groups working on repeated-query auditing of LLM brand recommendations and related behavioural testing [4, 5, 9]. We refer to this technique as the Dice Roll Method, drawing an analogy to the probabilistic process of rolling a die repeatedly to estimate its fairness: each “roll” (query) samples from the LLM’s output distribution, and aggregation across rolls yields increasingly stable estimates of the underlying distributional properties.“

“Despite its growing adoption, the Dice Roll Method lacks standardization across several critical dimensions. … No published work has established the statistical power of different iteration counts for detecting bias effects of various magnitudes, compared the reliability properties of different stability metrics, or provided guidance on cost-efficient experimental design for LLM auditing.

This paper addresses that gap by formalizing the Dice Roll Method as a standardized protocol for repeated-query auditing of LLM brand recommendations.“

Results

We’ll skip several technical sections and jump into the results.

In summary, the results show that AI visibility measurement becomes more reliable as identical prompts are repeated, but there is no universally correct iteration count. The appropriate number depends on the strength of the signal, the desired reliability, the model, and even the measurement method being used. The broader lesson for SEO/GEO professionals is to treat AI visibility as a statistical measurement problem rather than taking a handful of AI responses at face value.

There are also interesting findings within this section. One that caught my attention was the author found that model choice matters more than language choice in the multilingual data. Model-level variation was about 7.4 times greater than language-level variation, although there were model-language interactions. Gemini, for example, was less stable in Estonian and Finnish and consequently required more repetitions to reach comparable reliability.

Perhaps the headline finding — though it’s easy to misinterpret — is that more repetitions generally improve reliability, but with diminishing returns. In the authors’ internal data, precision improved quickly through roughly 7–10 runs, with about 80% of the study’s measured precision reached by 7 iterations, 90% by 10, and roughly 95% by 15. However, later external validation showed that these exact thresholds are dataset-dependent, reinforcing the paper’s broader “pilot-then-solve” recommendation rather than a universal run count.

Discussion

We’ll skip to the discussion section next.

This section turns the Dice Roll Method into a practical framework for AI visibility measurement — repeat prompts enough times to separate signal from stochastic variation, measure more than one dimension of visibility, and don’t assume a fixed number of runs will work everywhere.

The broader recommendation is to pilot the specific prompts, models, and conditions being studied, measure their variability, and then determine the repetition count needed for a reliable result.

In other words, “10 runs” or “15 runs” shouldn’t become an arbitrary industry benchmark. The real lesson is to measure the uncertainty in AI visibility rather than hiding it behind a single score.

Conclusion

“This paper formalizes the Dice Roll Method as a standardized protocol for repeated-query auditing of LLM brand recommendations, grounded in the generative mechanism of temperature-scaled nucleus sampling. …

The practical implications are operational. Researchers can now justify iteration counts through GLMM-based simulation-based power rather than normal-theory approximations, match target generalizability coefficients to audit-design budgets (iterations versus models), and pre-register drift diagnostics as a first-class component of every audit. Taken together, these protocol elements form a checklist that brand-recommendation audits can adopt. … By formalizing what has been an informal practice, we aim to strengthen the methodological foundations of a research community whose findings carry increasing policy and commercial significance.“

Knowing what we do now, why should SEO/GEO professionals care about “The Dice Roll Method: A Standardized Protocol for Repeated-Query Auditing of Large Language Model Brand Recommendations”?

In short, the broader lesson isn’t “run every prompt 15 times.” It’s that AI visibility should be measured with uncertainty in mind, using enough repeated observations to distinguish a real pattern from stochastic fluctuation.

This paper challenges a growing habit in AI visibility measurement of treating a small number of LLM responses as if they represent a stable ranking or recommendation outcome. The author shows that identical prompts can produce different brand recommendations across repeated runs, meaning a single response (or even a handful of responses) can be too noisy to support confident conclusions about brand visibility.

The Dice Roll Method offers a more rigorous alternative by treating AI visibility as a measurement problem under uncertainty. Rather than asking a prompt once and recording who appears, the protocol repeats identical queries, measures how results vary, and uses reliability, effect-size, semantic, and prompt-adjusted visibility metrics to determine whether observed patterns are stable in actuality.

For SEO/GEO, the broader implication is that AI visibility may be better understood as a distribution of outcomes than as a single ranking position. A brand may appear frequently, inconsistently, only for certain prompts, or with different competitors across repeated generations. That means practitioners should be cautious about reporting a single “share of voice” or visibility score without also considering variability and confidence.

The paper does provide rough iteration tiers (about 5–7 runs for exploratory work, 10–12 for confirmatory work, and 15–20 for more rigorous auditing) but the author cautions against turning those into universal rules.

The author’s external validation found that the statistical framework transferred better than the fixed iteration counts, supporting a pilot-then-solve approach: first measure the variability of the specific prompts, models, and conditions you are studying, then determine how many repetitions are needed.

To put it more succinctly, AI visibility should not be measured from a single response. Since LLM recommendations are stochastic, the same prompt can produce different brands across repeated runs. The Dice Roll Method provides a framework for measuring that variability more systematically, helping practitioners distinguish a real visibility pattern from random fluctuation.

But again, while AI visibility measurement becomes more reliable as identical prompts are repeated, there is no universally correct iteration count. The best number of runs depends on the strength of the signal, the desired reliability, the model, and the measurement method being used.

Caveat: The author is affiliated with Rankfor.AI, which implements the Dice Roll Method in a commercial AI brand-intelligence product. The five original brand-recommendation datasets reanalyzed in the paper also came from the author’s prior research. The paper includes independent external validation of its statistical methodology, although independent validation using another repeated-query brand-recommendation dataset remains unavailable.

Outro

I hope you’ve enjoyed this edition of Hamsterdam Research! 🐹

Feel free to comment below or contact me with your feedback.

Stay tuned for another new article, hopefully next week, or check out related research posts below.

Until next time, enjoy the vibes:

Thanks for reading. Happy optimizing!


Related research articles:

Examining New Research, “AI in Search Reduces Publisher Referrals Without Improving User Experience: Experimental Evidence,” for SEO/GEO Insights (Hamsterdam Research)

Examining New Research, “AI in Search Reduces Publisher Referrals Without Improving User Experience: Experimental Evidence,” for SEO/GEO Insights (Hamsterdam Research) By Ethan Lazuk Last updated: Welcome to another edition of Hamsterdam Research! 🐹 This is where we look at recent AI research papers to learn what they’re talking about and explore their hypothetical implications for…

Examining New Research, “Less can be More: Relieving RAG Bottlenecks via Evidence Frontloading and Pressure-Adaptive Budgeting,” for SEO/GEO Insights (Hamsterdam Research)

Examining New Research, “Less can be More: Relieving RAG Bottlenecks via Evidence Frontloading and Pressure-Adaptive Budgeting,” for SEO/GEO Insights (Hamsterdam Research) By Ethan Lazuk Last updated: Welcome to another edition of Hamsterdam Research! 🐹 This is where we look at recent AI research papers to learn what they’re talking about and explore their hypothetical implications…

Why Relevance Isn’t Enough for AI Search: The Importance of Answerability (Hamsterdam Research)

Why Relevance Isn’t Enough for AI Search: The Importance of Answerability (Hamsterdam Research) By Ethan Lazuk Last updated: Welcome to another edition of Hamsterdam Research! 🐹 This is where we look at recent AI research papers to learn what they’re talking about and explore their hypothetical implications for the future of search and SEO/GEO strategies.…

Editorial history:

Created by Ethan Lazuk on:

Last updated:

Need a hand with your SEO/GEO strategy?

I’m an independent SEO/GEO consultant based in New York City. Contact me for more information!

Leave a Reply

Discover more from Ethan Lazuk

Subscribe now to keep reading and get access to the full archive.

Continue reading

GDPR Cookie Consent with Real Cookie Banner