The Origins of Chunking Content (Viewed Through an SEO/GEO Lens)

Content chunking for AI search.

In the mythbusting section of its “Optimizing your website for generative AI features on Google Search” guidance, Google’s authors say:

“Chunking” content: There’s no requirement to break your content into tiny pieces for AI to better understand it. Google systems are able to understand the nuance of multiple topics on a page and show the relevant piece to users. However, sometimes shorter (or longer!) pages can work well depending on your audience and subject matter. There’s no ideal page length, and in the end, make pages for your audience, not just for generative AI search.

But Google isn’t saying “systems never segment text.” Its guidance just says publishers do not need to manually break pages into tiny pieces for Google’s AI features. That’s a much narrower claim, and it leaves room for passage-level retrieval inside the system.

What AI search practitioners now call “chunking” is the latest version of a much older IR problem where when a document contains many ideas, the system must decide whether to retrieve the entire document or identify the smaller part, the passage or chunk, that actually answers the question.

So should chunking be something you think about when creating content for SEO/GEO purposes? Let’s look at that question through a historical lens.

Where does chunking come from?

While the word “chunk” has roots in cognitive psychology — George Miller’s famous 1956 paper “The Magical Number Seven, Plus or Minus Two” discussed how humans overcome limits in short-term information processing through recoding information into larger meaningful units, or chunks — the most likely ancestor of AI-search chunking is passage retrieval in information retrieval. Now, your mind might go to Passage Indexing, which Google announced in 2020. However, the history goes back way before then.

In 1980, John O’Connor published research on “answer-passage retrieval,” which described systems that returned relevant passages rather than merely identifying entire documents. Answer-passage retrieval became a serious IR research area in the early 1990s, specifically with the work of Gerard Salton, James Allan and Chris Buckley. They wrote in 1993 that long documents often contain multiple topics and that retrieving entire documents can therefore be less useful than identifying the text excerpts most responsive to a user’s need.

Around the same time in the 1990s, Marti Hearst developed TextTiling. This technology automatically divided long documents into coherent multi-paragraph passages based on changes in vocabulary and subtopic. Hearst himself described the result as “passages, or subtopics,” while the suggested applications for TextTiling included information retrieval and summarization.

There’s plenty of other work in the 1990s worth mentioning, as well. James Callan’s 1994 work on passage-level evidence explored using paragraphs and text windows from long documents as evidence for document retrieval.

But so far we’ve talked all about information retrieval and chunking. There’s also a separate natural language processing (NLP) use of the term. Ramshaw and Marcus’s 1995 paper “Text Chunking using Transformation-Based Learning” dealt with identifying syntactic units such as base noun phrases. However, that’s called “shallow parsing,” which is different from today’s RAG-style document chunking, although the terminology overlaps.

Passage retrieval continued developing after that, but jump ahead to 2020 and we get Dense Passage Retrieval and RAG. That DPR/RAG period is where the connection to today’s AI search becomes especially direct.

That’s also when Google introduced Passage Indexing.

In the announcement for Passage Indexing, Google says “not just index web pages, but individual passages,” which caused confusion in the SEO community. Within days, Google clarified that it was not independently indexing passages as separate search documents. The page remained Google’s indexed object, only Google could now understand and use a particular passage as an additional signal when deciding whether that page was relevant to the query. Google consequently came to call the system “passage ranking,” rather than passage indexing.

And if you check Google’s ranking system documentation, it still mentions passage ranking: “Passage ranking is an AI system we use to identify individual sections or ‘passages’ of a web page to better understand how relevant a page is to a search.”

All that’s to say, Google unquestionably has modern systems that reason about sub-page units, which is to say passages or chunks.

So why then does Google say that chunking content for AI search is a myth?

If we parse the language carefully, Google isn’t saying that its retrieval systems never segment, identify, or evaluate passages. (Its own passage-ranking documentation would contradict that.) It’s merely saying that publishers don’t need to manually turn their content into tiny fragments as an AI optimization tactic.

But not everyone agrees that strategically chunking your content isn’t a worthwhile optimization tactic.

In “Google’s AI search guidance is naive and self-serving,” Mike King basically says that retrieval architecture matters to content optimization. His position is closer to “semantic chunking,” where sections should form coherent units that retain enough context to make sense when a retrieval system evaluates them independently.

I hold the same opinion. (Just check out any of my AI optimization articles, like the one about extractability.)

Plus, Google’s voice shouldn’t be the only one in the room, considering AI search platforms extend to multiple companies, like Microsoft and OpenAI.

For example, Microsoft’s writings support the underlying engineering point for chunking content. In May 2026, Bing published an explanation of how indexing changes when the consumer of search results is an AI grounding system rather than a human scanning ten blue links. Microsoft says the unit of value shifts from the document toward “groundable information,” i.e., discrete facts that can support an answer. It specifically says that “chunking/transformations must preserve meaning” and discusses the risks of breaking pages into retrievable chunks in ways that distort the original meaning.

In essence, Microsoft isn’t debating whether chunking exists. It’s discussing how to chunk without destroying meaning.

And now we have strong evidence that OpenAI uses chunks.

For example, OpenAI’s own retrieval infrastructure uses chunks. Its vector-store documentation, for instance, says the automatic chunking strategy uses a maximum 800-token chunk size with 400 tokens of overlap. Its File Search documentation is even more explicit, stating that the default settings are 800-token chunks, 400-token overlap, and up to 20 retrieved chunks added to the model’s context.

That said, we shouldn’t make the leap that ChatGPT Search indexes every webpage into the same 800-token chunks as OpenAI hasn’t publicly documented ChatGPT’s web-search pipeline that way. However, we do have some independent research on this topic.

In July and August 2026, RESONEO reverse-engineered ChatGPT’s web retrieval behavior by analyzing roughly 1,200 ChatGPT answers, 88,000 search results and 26,900 distinct pages. One interesting finding was that one observed retrieval pipeline represented pages using roughly 200 characters of stored snippet text, rather than reading the entire page every time. Those 200-character snippets were query-independent in the researchers’ observations, meaning the same stored representation could be returned for different searches involving the URL. That’s evidence that at least one observed ChatGPT retrieval pathway can operate using sub-page representations, or, arguably, chunks.

Takeaways

Passage retrieval has decades of IR history, going back to the 1980s with rich research in the 1990s. But it was in 2020 that Dense Passage Retrieval made passages explicit retrieval objects. RAG was then built around retrieved passages, and with passage ranking, we know that Google’s ranking systems understand passages.

Looking beyond Google, Microsoft openly discusses chunking for AI grounding, and OpenAI’s own retrieval products explicitly divide information into chunks.

What hasn’t been established is the claim made by some SEOs that publishers should format every page into some magic number (50 words, 100 words, 200 words, 300 tokens, etc.) because ChatGPT or Google will reward those blocks. I think it’s smarter to think about “semantic chunks,” or self-contained passages of content that can be extracted and understood in isolation.

Outro

Thanks for checking out my history on the origins of chunking content. You can find more related articles below.

And if you need help with your SEO/GEO strategies, feel free to contact me. I’m an independent consultant helping brands and agencies.

Until next time, enjoy the vibes:

Thanks for reading. Happy optimizing!

Related posts

Editorial history:

Created by Ethan Lazuk on:

Last updated:

Need a hand with your SEO/GEO strategy?

I’m an independent SEO/GEO consultant based in New York City. Contact me for more information!

Leave a Reply

Discover more from Ethan Lazuk

Subscribe now to keep reading and get access to the full archive.

Continue reading

GDPR Cookie Consent with Real Cookie Banner