Subscribe to Newsletter
EntitiesSemantic SEOKnowledge Graph

Entities over keywords: what Google's Knowledge Graph means for your SEO content

By 6 min read

An entity is, in the words of Google's patent, "a thing or concept that is singular, unique, well-defined and distinguishable." On May 16, 2012, Google introduced the Knowledge Graph with a now-famous phrase: "things, not strings." It marked the shift from matching keywords to understanding entities (people, places, concepts) and the relationships between them.

What is the difference between a string and a thing?

A keyword is a string of characters. An entity is a thing the search engine knows about, with attributes and connections. The patent's own example: there may be one node for the city "Philadelphia," one for the movie "Philadelphia," and one for the cream cheese brand "Philadelphia," each with its own unique identification reference. Same string, three entities.

The patent describes two mechanisms in the Knowledge Graph:

  • Disambiguation (one name, several entities): the city "New York" is told apart from the state "New York" because one is connected to the type "City" and the other to the type "State." Likewise, "Georgia" connected to "United States" is the U.S. state, while "Georgia" connected to "Asia" and "Eastern Europe" is the country.
  • Differentiation (several names, one entity): "George Washington," "Geo. Washington," "President Washington" and "President George Washington" may all point to a single entity node. California can carry aliases such as "CA," "Calif." and "Golden State."

For content, this means relevance is no longer only about repeating a phrase. It's about covering an entity and its relationships clearly so the page can be associated with the right entity, not a namesake.

What does the patent actually describe?

Google's patent US 10,235,423 B2, "Ranking search results based on entity metrics" (filed as a PCT application on December 12, 2012, granted on March 19, 2019, 27 claims), describes ranking search results obtained from a Knowledge Graph with several metrics determined at least in part from the Knowledge Graph, then combining them with weights that depend on the entity type of each result. The claims require that at least one of these metrics be "independent of the search query."

The patent names four illustrative metrics:

  1. Relatedness (R): based on the co-occurrence, on web pages, of the entity reference in the query with its entity type. The patent's example: the query "Empire State Building," of type "Skyscraper," where the co-occurrence of "Empire State Building" and "Skyscraper" in webpages may determine the metric. One formula given is N(E, REj) / (N(E) + N(REj) - N(E, REj)): the instances of both terms divided by the instances of either one, in a text corpus such as a set of webpages.
  2. Notable entity type (N): N = G / n, where G is a global popularity metric of the result (for example the likelihood that a random user selects it, the number of links to or from it, or visits) and n is the rank of its entity type in a list of notable types for the domain. The patent's example: a fiction novel with popularity 50 and type rank 1 gets 50; a short story with popularity 20 and type rank 8 gets 2.5.
  3. Contribution (C): based on critical reviews, fame rankings and other information, such as a restaurant rated 0 to 3 stars by a newspaper reviewer, best-seller lists for books, or reviews on Amazon, Yelp or IMDB. A weighting function makes the highest-reviewed contributions count most: "an actor's highest rated movies will have the greatest contribution to that actor's fame metric."
  4. Prize (P): based on awards and prizes, such as an Academy Award or Golden Globe for a movie, a Pulitzer or Man Booker prize for a book, a Nobel prize or military medal for a person. The most significant prizes weigh most.

The score is then a weighted sum: S = (a × R) + (b × N) + (c × C) + (d × P). The patent adds that this "particular calculation of a score is merely an example."

Why do weights change with the type of entity?

Because a metric does not say the same thing in every domain. The patent stores domain-specific weights, a domain being a group of entity types ("Books," "Film," "People," "Places"):

  • The system "may place the highest weight on the prize metric for entities with the 'Film' domain and may place the highest weight on a contribution metric for entities associated with the 'Book' domain."
  • The notable type metric may get more weight for "People" than for "Movie," because a movie's type (its genre) "does not necessarily give much information about its comparative relevance, while the entity type of a person, e.g., U.S. President, gives more information."
  • A weight can be 0 when a metric is not used in a domain, and even negative when a high value of a metric "is considered detrimental to the relevance" in that domain.

The weights can be determined experimentally, from system settings or from aggregated user selections. The resulting ranking may order a list of results, order image thumbnails, or decide which elements appear on a map.

How do you write for entities?

These are general content practices consistent with the patent's logic, not instructions taken from it:

  • Map the entity, not just the keyword. Before writing, list the sub-topics, attributes, and related entities a complete resource would cover. That map is your outline.
  • Use unambiguous language. Name things explicitly. Define the entity early. State its type ("X, a type of Y, used for Z"), since the patent disambiguates entities through their types and connections. Explicit names feed direct answers: Google answers "who" questions by counting entities named in the top results.
  • Add structured data. schema.org markup is a widely used way to state explicitly which entities a page is about. The patent itself does not mention schema.org.
  • Be consistent across the site. Reinforce the same entity associations through internal links and repeated, accurate references.

What does the entity metrics patent change for your SEO?

The patent means that for results tied to Knowledge Graph entities, relevance can be measured with signals about the entity itself, weighted by its type, and not only by the match with the query string. 4 consequences follow:

  1. Make the entity's type explicit. The type drives both disambiguation and the weights applied: a page that clearly says what kind of thing it is about is easier to place.
  2. Write the entity next to its type. Relatedness is computed from how often an entity and its type co-occur across web pages ("Empire State Building" with "Skyscraper"). Naming both together matches how the patent measures that link.
  3. Cite the signals that matter in your domain. Reviews and ratings for books, restaurants or products, awards for films, notable roles for people: the patent shows each domain can weight these differently.
  4. Don't expect one recipe for every topic. A signal that helps in one domain can count for nothing, or even count against, in another.

Our semantic SEO work maps each page to one main entity and states its type next to its name in the title, the opening paragraph and the structured data.

The patent describes what Google's system can do. It does not confirm that Google ranks results with these exact metrics or weights today.

Quick quiz

Did you get it?

Test what you just read.

Question 1 of 3

Which phrase did Google use to announce the Knowledge Graph in 2012?

Related articles

We read the patents so you don't have to.

Every week: 3 data-backed SEO insights, 1 myth busted, 1 pattern to steal. No fluff. No guru advice.

Free forever. Unsubscribe anytime. We respect your inbox. Privacy.