Does Language Change What AI Search Considers Authoritative? Lessons From 1,920 AI Overview Queries (Hamsterdam Research)

By Ethan Lazuk

Last updated:

Google and Baidu for health queries.

Welcome to another edition of Hamsterdam Research! 🐹

This is where we look at recent AI research papers to learn what they’re talking about and explore their hypothetical implications for the future of search and SEO/GEO strategies.

This time, we’ll look at research called, “Who Anchors AI Overviews in Health? Baidu, Google, and the Geography of Authority.”

It was submitted to arXiv on September 6, 2026, and its authors are Mingyue Zha and Ho-Chun Herbert Chang.

In short, why should SEO/GEO professionals care about this research?

This research suggests that AI search visibility isn’t universal, but rather it can change based on language, geography, and platform.

For SEO/GEO practitioners, the biggest implications are:

  • Authority is contextual: A source that appears authoritative in one country or language may not surface the same way elsewhere.
  • Language can materially affect citations: In this study, using a country’s official language increased locally sourced citations by roughly 3.5x to 13.5x.
  • Localization may matter for GEO more than simple translation: If AI systems retrieve different sources depending on language and country, localized content strategies could influence visibility.
  • Platform ecosystems matter: Google and Baidu both showed signs of routing users toward their own ecosystems, which affects which sources get exposure.
  • GEO measurement needs segmentation: A brand’s “AI visibility” score may look very different depending on the country, language, and search engine being tested.

The broader takeaway is that if AI search systems construct different information environments for different users, GEO strategy cannot assume one global set of authoritative sources.

One important caveat: this study is specifically about health queries, so we should treat the broader SEO/GEO implications as hypotheses rather than assume the same behavior applies equally across every industry

Let’s start by reviewing the paper’s abstract.

Here’s the abstract (with my highlights):

“Artificial intelligence is being rapidly incorporated into traditional search systems, yet scant work audits the information disparities across platforms, geography, and languages. We address this gap by comparing Google and Baidu’s AI Overview systems for health queries, and measure informational anchors that emerge. Auditing 1,920 health queries across 12 countries and 4 languages, we find that Google and Baidu exhibit vertical integration, routing users toward their own company platforms in AI Overviews rather than a diverse set of primary sources. Smaller, lower-localization countries receive fewer domestically sourced references for health queries. Issuing the same query in a country’s official language rather than English raises the share of locally sourced citations approximately 3.5- to 13.5-fold. Comparing queries across health topics of varying severity and controversy, including Traditional Chinese Medicine as an example, we also show that health disclaimers are multidimensional and vary across language and culture. We discuss how generative search influences access to health information, and the urgent need for culturally-aware oversight of these systems that influence critical health decisions.”

Next, let’s break down the paper’s key vocabulary terms:

  • Generative search: Search systems that use generative AI to synthesize an answer rather than simply returning a ranked list of links. The paper examines Google and Baidu as examples.
  • AI Overviews (AIOs): AI-generated summaries shown within search results that synthesize information from multiple sources and provide citations. The paper treats these as a distinct information layer above traditional organic results.
  • Informational anchors: The sources, institutions, platforms, or other factors that shape where an AI Overview gets its information. The authors identify corporate ownership, geography, and language as three major anchors.
  • Vertical integration: When a search platform disproportionately directs users toward content or properties within its own corporate ecosystem. The study finds this behavior in both Google and Baidu’s AI Overview citations.
  • Source localization: The extent to which an AI-generated answer cites sources from the same country as the user or query. The researchers measure this partly through country-specific domains.
  • Localized citations: Citations pointing to domestically relevant sources rather than U.S. or global institutions. The study found that changing the query from English to a country’s native language could dramatically increase these local citations.
  • Query language: The language in which a search is performed. In this study, query language affected which geographic sources appeared even when the user’s country stayed the same.
  • Source diversity: The variety of different publishers, institutions, and information sources represented in an AI-generated response. The researchers found that both Google and Baidu often concentrated citations among a relatively small group of intermediary platforms.
  • AI Overview return rate: The percentage of queries for which a search engine actually generates an AI Overview. The study found that return rates varied depending on platform and topic sensitivity.
  • Zero-shot labeling: Using a language model to classify or score content without first training it on labeled examples for that specific task. The researchers used this method to analyze things like support, skepticism, medical referrals, and safety warnings.

Bien. Let’s take a deeper look at the research paper’s contents now.

If you want to follow along, you can grab a PDF or the HTML version on arXiv.

The paper has five main sections.

We’ll focus on the most relevant ones below.

Introduction

“Search is arguably one of the most consequential applications of generative artificial intelligence (AI),” the authors write. “Every day, Google processes approximately sixteen billion search queries, over one billion of which concern healthcare.”

The authors next explain Google’s global dominance, except for in China, hence Baidu:

“From 2024 to 2025, overall exposure to Google AI Overviews (AIO) expanded from 7 to 229 countries, due to Google’s dominance as a global search engine [2]. However, one country remains excluded from Google’s dynamics: China. China’s Great Firewall regulates and censors domestic internet, including Google and all of its related products [7]. Instead, Chinese netizens use Baidu, a major Chinese company often called the “Google of China”, which offers a range of online services, mobile apps, and AI technologies. In terms of market capture, Baidu commands 63.97% of the search market for 1.1 billion Chinese internet users [8]. However, Baidu accounts for less than 1% of the total search market outside of China [9]. This has made Baidu region-specific and localized for Chinese users.”

They next talk about AI Overviews, which promise “to lower the cognitive burden of a search but narrows a user’s exposure to source diversity [and] Because of this placement, AI Overviews carry an out-sized influence over public knowledge, opinions, and people’s behavioral intentions.”

Then they drop some stats about user behavior: “Google’s AI Overviews reduce outbound organic clicks by roughly 39.8% and raise the rate of zero-click searches by 34.5% when the feature appears, with actual visits to an AI Overview source in only about 1% of cases.”

Consequently, “biases carried by the model from its training data or its retrieval process may propagate to billions of users.”

Next they talk about localization and personalization: “Across studies, IP address and login state are consistently the strongest predictors of how much a result set is personalized,” they write. “Personalization in search may be helpful for filtering for relevancy, but may also lead to informational disparities.”

This partly impacts the lack of comparability of results among search engines: “Audits of Google, Bing, DuckDuckGo, Yahoo, Baidu, and Yandex found substantial cross-engine differences, with the overlap between Google and Bing always under 32% for a given query.”

They explain that most of the research focus is on Google, but that leaves out China. “China, which accounts for 17% of the world’s population,” they write, “has a distinct information ecosystem and arguably the largest information bubble in existence.” That said, “studies comparing Google and Baidu are sparse.”

One study comparing Google and Baidu “compared 6,320 query results across the two engines and found only 6.8% overlap and little ranking similarity.”

“Chinese search engines also rarely direct users beyond national borders and disproportionately favor the country’s own content,” they write, “with 32% of Baidu’s results linking to Baidu’s own properties compared with 8% for Google.”

Next they talk about language’s impact on search results: “Language can also shape search results independently of geography or platform,” they write. “For LLMs, the use of different languages can surface different levels of bias and accuracy.”

“As of early 2026, Google’s AI Overview appears for 64.7% of question-form queries overall,” they write. “In this study, we investigate how AI Overviews diverge across health-related queries.”

Results

“Using SerpApi’s search APIs, we collected 1,920 responses to health queries across 12 countries, 4 languages, and two platforms, Google and Baidu,” they write.

“First, we examined the ten most-cited domains for Google’s and Baidu’s health AI Overviews. Citations indicate the authorities and sources that AI Overview systems draw on to synthesize their answers, at times quoting directly from the referenced material [35, 29]. They thus serve as a proxy for the types of institutions and information channels that AI Overviews rely on for health guidance. Both systems concentrated their citations on a relatively small set of intermediary platforms rather than on a diverse range of primary sources (Figure 1).”

Figure 1 from the paper.

“Similar to Google, Baidu concentrated citations within its own corporate ecosystem,” they write.

Thus, “Both systems exhibit vertical integration in AI Overview references through different intermediaries [which] is consistent with prior findings that search engines favor properties tied to their own commercial interests.”

Interestingly, “Mayo Clinic, despite being the most-cited institutional domain in US Google search, is not in the top 10 of YouTube channels cited, while non-institutional channels such as the Infographics Show and Doctor O’Donovan are.”

“This suggests that the authority of the cited source may depend on channels surfaced by YouTube’s ranking mechanisms rather than on health-specific editorial review,” they write. “As a result, reliance on YouTube may introduce a vulnerability in the credibility of AI health overviews, since the quality and authority of cited content can vary substantially even when all of them appear under the YouTube domain.”

On the topic of localization, their findings “suggest that AI Overviews frequently prioritize established U.S. and international health institutions over local sources in smaller markets, even when responding to queries submitted from countries with their own established national health authorities.”

However, they found that query language played a big role in determining search results:

“… localized sourcing was consistently low when queries were issued in English, accounting for 16.6% of citations in Japan, 11.5% in Singapore, 4.6% in Malaysia, and 2.7% in South Korea. In contrast, issuing the same queries in the country’s native language substantially increased the share of local sources, reaching 57.5% in Japan, 42.1% in Malaysia, 40.2% in Singapore, and 36.4% in South Korea. Across countries, this represents an approximately 3.5- to 13.5-fold increase in localization, ranging from a 3.5-fold increase in Japan and Singapore to a 9-fold increase in Malaysia and a 13.5-fold increase in South Korea. This indicates that query language is an important determinant of whether locally relevant sources show up in AI Overview citations.“

“Consequently,” they write, “users who search for health information in English, including local users who do so out of habit, educational background, or professional necessity, are less likely to receive information from the local country’s health information ecosystem, regardless of where the query is submitted geographically.”

They also found that “Certain conditions receive more localized citations than others [which] may be because these conditions have more established, localized institutional resources and support.”

They also looked at the AI Overviews return rate based on query type:

“Google’s return rate declined steadily as topic sensitivity increased, from roughly 90% of queries for common conditions down to roughly 42% for TCM queries. Baidu’s return rate followed a less consistent pattern: it matched Google’s high rate for common conditions, dipped to its lowest point for legally variable topics, and then rose again for TCM queries, where it exceeded Google’s return rate.”

They also found that “Averaged per health category, Google’s AI Overviews drew on more source references than Baidu’s overviews in every category.” This suggests that “Baidu is often as likely, or more likely, than Google to generate an AI Overview for a given query, but cites substantially fewer sources when it does so.”

Discussion

I’ll summarize the key takeaways from this section in my own words, with an SEO/GEO angle:

Platform ownership, geography, and language strongly shape AI Overview sourcing: The paper’s core argument is that generative search is not sourcing information in a neutral, uniform way.

Platforms tend to favor their own ecosystems: Google cited YouTube more than the CDC, WHO, or Mayo Clinic, while Baidu frequently cited Baidu-owned properties. That suggests platform ownership can influence which sources receive visibility.

“Authority” can depend on the platform: Google’s reliance on YouTube means a hospital channel and a general-interest creator can both appear under the same YouTube domain, so domain-level authority does not necessarily equal content-level authority.

Larger, established markets were more likely to receive localized sources: Australia, the UK, the US, and Canada saw more domestic citations than countries such as India, Nigeria, Norway, and Singapore.

Topic matters too: Health topics with strong local infrastructure, such as HIV, diabetes, depression, and long COVID, received more localized citations, while more politically, legally, or culturally contested topics received fewer.

The authors suggest two possible reasons for weak localization: Local authoritative sites may have less digital visibility, or algorithms may intentionally lean on globally recognized sources for sensitive topics and smaller markets. The study cannot tell which explanation is responsible.

Language can change the source ecosystem even when geography stays the same: Switching from English to a country’s official language increased the share of citations from local sources, suggesting multilingual search can improve local relevance.

Different platforms handle contested topics differently: For Traditional Chinese Medicine, Baidu was more supportive but also gave stronger safety warnings, while Google was less likely to generate an AI Overview at all.

“Caution” in generative search is more than whether an answer contains a disclaimer: The authors argue caution includes tone, safety warnings, citation quantity and quality, and even whether the system chooses to answer.

The findings are a snapshot, not a permanent rule: Google and Baidu are fast-changing systems, and the study mainly analyzes citations at the domain level, so it cannot fully judge the accuracy or quality of individual cited pages.

Conclusion

“Our work offers several new insights into how generative search impacts access to health information. We show that the sources cited by AI Overviews vary by platform, country, language, and health topic. Healthcare is not a monolithic science. How healthcare is organized and practiced is often reflective of the history and values of the society delivering it. For instance, Taiwan’s national health insurance formally covers Chinese medicine alongside Western medicine, a coexistence dating to 1958 [39]. Physicians in China and Taiwan often combine biomedical and traditional approaches rather than treating them as mutually exclusive systems [40]. Medical anthropologists have long argued that even Western “biomedicine” itself is a culturally situated system, not a value-free default against which other traditions are measured [41, 42].

As AI Overviews continue to expand, we hope this study contributes to a more informed conversation about how generative search should be evaluated and governed as a health information intermediary. Decisions about who can be a health authority should be made more deliberately rather than left to the incidental effects of platform ownership, query language, and retrieval pipelines. Making sure generative responses are reflective of relevant and local health contexts is important for everyone who uses these systems, and it remains particularly important for languages and populations that have been historically under-served in global information systems.“

Knowing what we do now, why should SEO/GEO professionals care about “Who Anchors AI Overviews in Health? Baidu, Google, and the Geography of Authority”?

SEO/GEO professionals should care about this research because it shows that visibility in generative search is not determined by content alone: platform, geography, language, and topic can all influence which sources are surfaced and cited.

Google and Baidu both showed signs of favoring sources within their own ecosystems, suggesting that platform ownership can shape the information environment users see. The study also found that local-source visibility varied widely by country, with larger and more established markets receiving more domestic citations than several smaller markets.

Perhaps the most important GEO finding is the effect of language. When researchers kept the country the same but changed queries from English to the country’s official language, the share of locally sourced citations rose significantly. That suggests multilingual GEO is not simply about translating content but rather the language of the query itself may change which sources an AI system considers relevant or authoritative.

The research also complicates the idea of “authority.” On Google, YouTube was the most-cited domain for health queries, ahead of organizations such as the CDC, WHO, and Mayo Clinic. But a YouTube citation could point to a major hospital or to a general-interest creator, meaning domain-level authority does not necessarily tell us much about the authority of the individual source.

For SEO/GEO professionals, the broader lesson is that AI visibility is contextual. A brand or publisher may be highly visible for one language, market, platform, or topic and far less visible for another. Measuring and optimizing GEO therefore requires thinking beyond a single universal visibility score or citation strategy.

Caveat: This study examines health queries specifically, so we should not assume the exact same patterns apply to every industry. But it provides strong evidence that GEO strategies and measurement should account for platform, localization, language, and the type of query being asked.

Outro

I hope you’ve enjoyed this edition of Hamsterdam Research! 🐹

Feel free to comment below or contact me with your feedback.

Stay tuned for another new article, hopefully next week, or check out related research posts below.

Until next time, enjoy the vibes:

Thanks for reading. Happy optimizing!


Related research articles:

Examining New Research, “Less can be More: Relieving RAG Bottlenecks via Evidence Frontloading and Pressure-Adaptive Budgeting,” for SEO/GEO Insights (Hamsterdam Research)

Examining New Research, “Less can be More: Relieving RAG Bottlenecks via Evidence Frontloading and Pressure-Adaptive Budgeting,” for SEO/GEO Insights (Hamsterdam Research) By Ethan Lazuk Last updated: Welcome to another edition of Hamsterdam Research! 🐹 This is where we look at recent AI research papers to learn what they’re talking about and explore their hypothetical implications…

Why Relevance Isn’t Enough for AI Search: The Importance of Answerability (Hamsterdam Research)

Why Relevance Isn’t Enough for AI Search: The Importance of Answerability (Hamsterdam Research) By Ethan Lazuk Last updated: Welcome to another edition of Hamsterdam Research! 🐹 This is where we look at recent AI research papers to learn what they’re talking about and explore their hypothetical implications for the future of search and SEO/GEO strategies.…

How Many Times Should You Ask an AI Assistant the Same Question? Research Into Measuring AI Visibility (Hamsterdam Research)

How Many Times Should You Ask an AI Assistant the Same Question? Research Into Measuring AI Visibility (Hamsterdam Research) By Ethan Lazuk Last updated: Welcome to another edition of Hamsterdam Research! 🐹 This is where we look at recent AI research papers to learn what they’re talking about and explore their hypothetical implications for the…

Editorial history:

Created by Ethan Lazuk on:

Last updated:

Need a hand with your SEO/GEO strategy?

I’m an independent SEO/GEO consultant based in New York City. Contact me for more information!

Leave a Reply

Discover more from Ethan Lazuk

Subscribe now to keep reading and get access to the full archive.

Continue reading

GDPR Cookie Consent with Real Cookie Banner