Subscribe to Newsletter
EntitiesPatentsDirect answers

Entity answers and SEO: how Google answers "who" questions by counting entities in the top results

By 8 min read

An entity answer is a direct answer to a question, chosen among the entities of the expected type that the top search results mention most and most topically. Patent US 9,477,759 B2, "Question answering using entity references in unstructured data," describes how Google selects it. Google filed it on March 15, 2013, and the patent was granted on October 25, 2016.

What does patent US 9,477,759 describe?

Patent US 9,477,759 describes a question answering system that reads the entities mentioned in the top results of a query and returns the best-ranked one as the answer. Its 2 inventors are Dvir Keysar and Tomer Shmiel, who also co-signed the related questions patent. The patent holds 21 claims and 8 drawing sheets.

The patent starts from a limit of earlier systems: answers "have been determined based on previously user-answered questions and manually generated databases." Here, answers "may be identified automatically based on unstructured content of a network such as the Internet," and the system can "take advantage of search result ranking techniques."

The abstract defines the method: a query "associated at least in part with a type of entity" is received, search results are generated, "previously generated data" with "one or more entity references in the at least one search result corresponding to the type of entity" is retrieved, the entity references "are ranked," and "an answer to the query is provided based at least in part on the entity result."

What is an entity, and what is a type of entity?

An entity is "a thing or concept that is singular, unique, well-defined and distinguishable," such as a person, a place, an item or an idea. A type of entity is a category that the answer must belong to: "types may include persons, locations, movies, musicians, animals, and so on." The patent focuses on "who" questions, whose answers are of the person type, and states that the same technique can answer "what" or "where" questions.

A person is defined broadly: "an individual, a group of people, a company, a legal entity, a team." The patent's examples of "who" queries include [Who was the first to climb Mount Everest?] and [Who won the 1997 World Series]. A "who" query can also be implicit, without a question mark or an interrogative: the query [first to climb Mount Everest] "is interpreted as a 'who' query."

How does Google find the answer?

Google finds the answer in 5 steps:

  1. Receive a query associated with a type of entity, such as a "who" question.
  2. Retrieve the top results, "for example, the top ten ordered search results."
  3. Retrieve the entity references of each result, a list usually built offline in advance (it can also be built at the time of the search) with, for each entity, an identifier, "a frequency of occurrence," "the location on the page where the entity reference occurs" and metadata such as freshness.
  4. Aggregate and rank the entity references across the results.
  5. Return the top entity as the answer, as text or as a natural language sentence, optionally with "a picture" or "a link to an encyclopedia entry," displayed "along with, or in place of, the search results."

How does the "king of Spain" example work?

The patent's drawing shows the query [Who is the king of Spain?] with 4 results, of which the system uses the top three. The first, titled "Monarchy of Spain," mentions Juan Carlos 2 times (once in the snippet, once elsewhere in the result), Philip II, "a 16th century king of Spain," 1 time, and Mike Jones, "the author of a book on Spanish Kings," 1 time. The second mentions Juan Carlos 5 times, Sophia of Greece 3 times and Bob Smith 1 time. The third mentions Philip II 2 times, Isabella 2 times and Charles V 1 time.

The summation counts "7 instances of [Juan Carlos I]" in the top three results, against 3 for Philip II and 3 for Sophia of Greece, and the system answers "The King of Spain is Juan Carlos I." The drawing shows the names in the snippets, but the patent specifies that entity references "may appear in any suitable content of the webpage," including its unstructured text.

The example dates from the filing in 2013. Juan Carlos I abdicated in 2014: the same counting today would rely on pages updated since. Freshness, which the patent lists among its signals, is what would let the answer follow such a change.

Which signals rank the entity references?

2 main signals rank the entity references, combined in "a weighted combination of ranking signals," where the weighting can include "the ranking of the search results or other search quality metrics":

  1. Frequency of occurrence: "the total number of times an entity reference appears in a document," possibly normalized "by the length of the document."
  2. Topicality score: it can include "freshness, the age of the document, the number of links to and/or from the document, the number of selections of that document in previous search results, a strength of the relationship between the document and the query."

Topicality depends on the page as well as the entity. The patent's examples: "George Washington" has "a higher topicality score on a history webpage than on a current news webpage," and "Barack Obama" a higher score "on a politics website than on a law school website." The ranking and selection can also rely on "a quality score, a freshness score, a relevance score."

What do the claims of patent US 9,477,759 protect?

The 21 claims are grouped under 3 independent claims (a method, a system and a computer-readable medium), and they are more precise than the description:

  1. The search results are ranked "based on relevance to the query," and only the subset "above a first predetermined ranking threshold" is counted.
  2. For each entity reference, the system computes "a weighted sum of the frequencies of occurrence" in each result of that subset, a sum that normalizes the frequency for each set of previously generated data, then ranks the entities by that sum.
  3. In case of a tie, the dependent claims add the results "below the first predetermined ranking threshold and above a second predetermined threshold," compute a ranking signal from them and rerank the entities.
  4. Other dependent claims normalize the frequency by "a length of the respective previously generated data," compute a topicality score "based on the number of links to and from" that data, and state that the data can be unstructured.

How does Google make sure the answer is reliable?

Google builds confidence by adding results until one answer stands out. "Processing only the top two search results may yield a tie or near-tie," so the system "may successively add previously generated data from lower ordered search results until a particular answer is significantly more common in the results than the others." The number of results processed can depend on "system design, user preferences, the type of query," system speed, previous question answering or "the quality of an identified answer."

Entity references are identified "by comparing words and phrases in the content with known entity references," such as a database of names. Unknown names can be found by clustering: "a person's first and last name that appear together repeatedly in unstructured text may be identified as an entity reference." The other names in the content disambiguate a reference: [George Washington] next to [Martha Washington] is the U.S. President, while [George Washington] next to [University] and [Washington D.C.] is George Washington University.

The patent also describes a knowledge graph in which entity references can be stored and which separates entities sharing a name. Its examples: the city, the movie and the cream cheese brand named "Philadelphia" are distinct nodes with their own identifier, and the city of New York is distinguished from the state because one connects to the entity type [City] and the other to [State].

What does entity answering change for your SEO?

Entity answering means that the answer Google gives is the entity that the top pages name most often and most centrally, so consistent, clear naming across your content feeds the answer. 5 consequences follow:

  1. Rank in the top results. The answer is built from the entities of the top results, "for example, the top ten," and the claims only count the results above a ranking threshold.
  2. Name the answer explicitly. Frequency counts, and references are found by matching names: state the entity's name in the text rather than relying on pronouns.
  3. Make your page about the topic of the question. Topicality rises when the entity is central to the page, as George Washington on a history page.
  4. Keep factual pages up to date. Freshness is part of topicality, and the king of Spain example shows how fast an answer can change.
  5. Disambiguate your entities. Use the full name and mention related entities (place, dates, roles, related people), as Martha Washington identifies the President, so the system links your mention to the right entity, as described in the Knowledge Graph patent.

The patent describes what Google's system can do. It does not confirm how Google selects direct answers today.

Quick quiz

Did you get it?

Test what you just read.

Question 1 of 4

Where does the system look for the answer to a "who" question?

Related articles

  • Featured snippets
    US 9,940,367 B1
    Featured snippetsPatentsAnswer passages

    Answer passage scoring for SEO: how Google picks the passage that answers a question

    Patent US 9,940,367 describes how Google builds candidate answer passages from the top results and scores them with query terms, expected answer terms and 6 query-independent signals. The full mechanism and how to write passages that get picked.

We read the patents so you don't have to.

Every week: 3 data-backed SEO insights, 1 myth busted, 1 pattern to steal. No fluff. No guru advice.

Free forever. Unsubscribe anytime. We respect your inbox. Privacy.