The Origins of Entities: How Knowledge Graphs and Semantic Search Shaped Search Engines

Before we understand the origins of entities, we must understand the problem: words can mean different things.
Jaguar can refer to an animal, a car brand, or an NFL team player. Or where I live can go by NYC, New York City, or the Big Apple.
Entities help solve for that. But while entities, knowledge graphs, and semantic search are all related, each describes something different.
Remember this terminology:
| Concept | Explanation |
|---|---|
| Entity | An identifiable thing, such as a person, organization, place, product, or concept. |
| Entity mention | The words used to refer to that thing in a particular passage. |
| Entity linking | Connecting a mention to the appropriate entry in a knowledge base. |
| Knowledge graph | A structured representation of entities, their properties, and their relationships. |
| Semantic search | Search that uses meaning and context to interpret queries and find relevant information. |
In short, entities give machines a way to represent knowledge.
Now let’s talk about entities’ foundations.
In the days before Google, knowledge representation, semantic networks, and database modeling were all earlier attempts to describe things and their relationships.
Then came the idea of the Semantic Web, the ambition to make information on the web usable by machines through explicit descriptions and connections.
RDF (Resource Description Framework) was already a W3C recommendation in 1999, long before Google’s Knowledge Graph appeared. It was based on simple subject–predicate–object statements, or what we call semantic triples, like Ethan Lazuk–works as a–consultant.
The Semantic Web is a broader vision for machine-readable and connected data. Meanwhile, semantic search is a search capability. Those two are distinct.
But where did the knowledge for the Semantic Web come from?
Some sources included Wikipedia, DBpedia, Freebase, and Wikidata, which served as encyclopedia articles and structured databases for factual information.
Metaweb and Freebase deserve particular attention, in that regard.
Google acquired Metaweb in July 2010 and connected its acquisition to understanding real-world entities and their relationships. Google’s announcement even discussed answering questions with several constraints, such as finding actors of a certain age who had won an Oscar.
Meanwhile, Wikidata is useful for demonstrating how an entity can have an identifier as well as several language labels, statements, and references.
In 2011, search engines agreed on a vocabulary: schema.
Schema.org launched jointly by Google, Bing, and Yahoo in June 2011.
This meant publishers could describe people, organizations, products, events, and miscellaneous things using a shared vocabulary.
But remember, while Schema.org provides the vocabulary for describing information, a knowledge graph is what organizes that information. Adding schema markup by itself does not establish that a search engine accepts every claim.
Google introduced its Knowledge Graph on May 16, 2012.
This was a central historical moment when entity information became highly visible in mainstream search.
Google’s Knowledge Graph had three practical benefits, including distinguishing between meanings, summarizing facts, and helping people explore related things. Their announcement also identified sources, including Freebase, Wikipedia, and the CIA World Factbook.
We’ve talked before on this blog about knowledge panels. That’s a presentation feature, meanwhile the Knowledge Graph is the underlying information resource.
Also in 2012, Bing had a parallel story with Snapshot and Satori.
Bing’s Snapshot appeared in June 2012, displaying information about entities directly in search results. In March 2013, Microsoft explained Satori, the underlying technology for representing entities and relationships, and described its expansion across people, places, and things.
The examples in the Satori expansion announcement included facts about landmarks, professional information, relationships between people, and answers to conversational questions.
Between Bing’s Satori and Google’s Knowledge Graph, we have valuable evidence that entity understanding was an industry-wide direction as of 2012.
In 2013, another milestone for entities came with Google’s Hummingbird update.
The Knowledge Graph and Hummingbird are separate entities, so to speak. The Knowledge Graph organizes information about entities, while Hummingbird was a major improvement to Google’s overall ranking systems for interpreting complex queries.
Google dates Hummingbird to August 2013 but now lists it under historical/retired systems. That means we should avoid describing present-day Google as simply running the original Hummingbird algorithm.
The next few years brought meaning beyond explicit graphs with RankBrain, neural matching, and BERT.
This is a second technological thread from the Knowledge Graph. Now we’re talking about systems that learn representations of language and concepts.
Google’s own retrospective provides a useful sequence:
- RankBrain, 2015: relating words to concepts to improve ranking.
- Neural matching, 2018: matching the concepts represented in queries and pages.
- BERT, 2019: interpreting combinations of words in context.
These developments show why semantic search is broader than looking up entities in a graph. They also demonstrate that new systems can work alongside existing ones.
This all brings us to where we are today, with entities, retrieval, and generated answers.
Today we have AI Overviews, AI Mode, and Bing Copilot which use retrieval and grounding to find supporting information and use it to construct an answer.
Google describes how its AI search features can issue related searches across subtopics (fanout queries). Meanwhile, Microsoft’s 2026 discussion explains how an index increasingly supports evidence gathering for generated answers, with attention to attribution, freshness, and preserving factual meaning.
Microsoft’s GraphRAG provides a concrete example of graphs and language models working together. Microsoft’s research uses an LLM to extract an entity graph from documents and create summaries of connected communities. Note that this is a documented research approach and doesn’t establish that Google or Bing uses that architecture for every answer.
What does this history of entities mean for SEO/GEO?
In short, publishers benefit from making identities, relationships, and supporting claims clear.
Useful questions to ask yourself include:
- Can someone clearly identify the organization, author, product, or service?
- Are relationships explained accurately and consistently?
- Do sources support the claims being made?
- Does structured data agree with the visible content?
- Does the page answer the actual question, including its context and constraints?
Entity mentions are not a checklist that guarantees rankings or AI citations. Google’s guidance continues to emphasize established SEO practices and says its AI features don’t require special Schema.org markup.
Rather, the thread that carries throughout entity history is identity, relationships, context, and evidence. That’s how AI search systems might understand a book, its author, its adaptations, and a reader’s increasingly complex questions about it.
Outro
Thanks for checking out my history on the origins of entities. You can find more entity-related articles below.
Or, if you need help with your SEO/GEO strategies, feel free to contact me. I’m an independent consultant helping brands and agencies.
Until next time, enjoy the vibes:
Thanks for reading. Happy optimizing!
Related posts
Entity Optimization for AI Search: Reducing Ambiguity and Increasing Retrievability
Entity Optimization for AI Search: Reducing Ambiguity and Increasing Retrievability Entity optimization for AI search isn’t a checklist of SEO/GEO tactics; it’s a holistic process…
How Do You Build Entity Authority So AI Assistants Stop Confusing Your Brand with Similar Names? A Tales from the Query Post
How Do You Build Entity Authority So AI Assistants Stop Confusing Your Brand with Similar Names? A Tales from the Query Post This article will…
Query Fanout as an Entity Optimization Strategy for AI Search
Query Fanout as an Entity Optimization Strategy for AI Search I’ve written before about how entity optimization for AI search isn’t a checklist of tactics…
Leave a Reply