Subscribe to Newsletter
Featured snippetsPatentsAnswer passages

Answer passage scoring for SEO: how Google picks the passage that answers a question

By 10 min read

Answer passage scoring is the method Google uses to rate the passages that compete to answer a question, by combining how well they match the query, how well they match the likely answer, and signals of passage quality. Patent US 9,940,367 B1, "Scoring candidate answer passages," describes it. Google filed it on August 12, 2015, from a provisional application of August 13, 2014, and the patent was granted on April 10, 2018.

What does patent US 9,940,367 describe?

Patent US 9,940,367 describes a system that generates candidate answer passages from the top-ranked resources for a question query, scores each one, and selects the best for an answer box. Its 8 inventors are Steven D. Baker, Srinivasan Venkatachary, Robert Andrew Brennan, Per Bjornsson, Yi Liu, Hadar Shemtov, Massimiliano Ciaramita and Ioannis Tsochantaridis. The patent holds 20 claims and 12 drawing sheets.

The answer passage is shown "separate and distinct from the search results, e.g., as in an 'answer box'." The patent does not use the words "featured snippet," but it describes this kind of answer.

Where do candidate passages come from?

Candidate passages come from the top N resources returned for the query. The patent notes that "the value of N may vary," and that in some implementations N is "the same number as the number of search results returned on the first page of search results." It adds that "a larger set of resources can also be used," so the top N is the described default, not an absolute limit.

Each resource is cut into passage units: "a complete sentence, a portion of a sentence, a header, or content of structured data, such as a list entry or a cell value." The system assembles candidates from units that pass a set of selection criteria.

Which content can enter a candidate passage?

Content enters a candidate passage only when it passes criteria that differ for running text and for structured content. For running text, the patent gives 7 examples, which it calls "illustrative":

  1. Complete sentences: a partial sentence is dropped or extended until a full sentence is detected.
  2. A minimum number of words: short units are dropped or extended.
  3. Visibility: text "rendered so that it is invisible to a user" is excluded.
  4. No boilerplate: text detected as boilerplate is excluded.
  5. Alignment: content that is not contiguous with the passage is excluded.
  6. One section only: text subordinate to different headings does not mix in the same passage, and a heading can only be the first sentence.
  7. No image captions: a caption cannot combine with other units.

The patent also mentions a maximum passage size and the "exclusion of anchor text" as further criteria.

Which rules apply to lists and tables?

Lists and tables follow structured content criteria that aim to keep the answer complete and coherent. The patent describes them as follows:

  1. Incremental list generation: the system takes one unit from each list element before taking a second one, so that "a complete list is more likely to be generated." In the patent's thermometer example, the answer keeps only the first sentence of each list item, because a maximum size was reached: "in short lists, the second sentence of a multi-sentence list element is less informative than the first sentence."
  2. All steps in a step list: when preferential ordering terms such as "steps," "first," "second" are detected, "all steps are included." If including all steps exceeds the maximum passage size, the candidate is discarded in some implementations, while in others the size limit is ignored for that passage.
  3. Superlative ordering: for a superlative query such as [longest bridges in the world], the system selects rows in descending order of the attribute, for example the rows for the 3 longest bridges out of a table of 100.
  4. Informational queries: for [nutritional information for Brand X breakfast cereal], the system can select "the entire table."
  5. Entity attribute queries: for [calcium nutrition information for Brand X breakfast cereal], it selects only the values that describe calcium.
  6. Key value pairs: each unit must include a complete key value pair, never a key without its value.

Mixing prose and structure follows 2 more rules. Once structured content is in a passage, "subsequent unstructured content cannot be added." And the sentence just before a table is checked for an enumerating reference: in the baggage fees example, it begins with "These," so "only the sentence is included" before the table; otherwise, "two or more sentences preceding the structured content are included."

How does Google score a passage against the query?

Google gives each passage a query dependent score that combines 2 match scores:

  1. Query term match score: the similarity between the query terms and the passage. In some implementations it is "proportional to a number of instances of matches of query terms to terms of the candidate answer passage," with query terms "weighted, e.g., by term frequency/inverse document frequency (TF/IDF) values."
  2. Answer term match score: the similarity between the passage and the terms likely to appear in the answer.

The 2 scores are summed, multiplied or "combined in other appropriate ways."

What are answer terms?

Answer terms are the terms that the top-ranked results use to answer the question. The patent explains why they are needed: question queries "do not describe what the user is looking for, as the answer is unknown to the user." In some implementations, the likely answer terms are "derived from the top N ranked resources returned for the query." The system builds the score in 5 steps:

  1. List the terms of the top-ranked resources in a term vector; stop words can be omitted.
  2. Weight each term by "a number of resources in the top-ranked subset of resource in which the term occurs multiplied by an inverse document frequency (IDF) value." The IDF can come from a large corpus of documents or from the top N documents.
  3. Count each term in the candidate passage.
  4. Multiply the weight by the count.
  5. Combine the multiplied weights, for example by summing them, into the answer term match score.

The patent's example uses the term "apogee" with a weight of 0.04: a passage that contains it twice earns 0.08, and a passage that contains it 3 times earns 0.12. With this formula, a term that appears in many top results and has a high IDF carries the most weight: it is the vocabulary of the answer.

How does the expected entity type change the score?

The answer term match score drops when the passage names no entity of the type the answer requires. The system determines an entity type for the answer, either from the terms with the highest scores or from the query: for [who is the fastest man], the entity type is "man." It then identifies the entities described in each passage. In the patent's example, a passage about Olympic sprinters and the 100 meter dash loses points, because "the term 'sprinter' is gender neutral." The score can be binary (1 if a term of the right type is present, 0 if not) or a likelihood that the correct term is in the passage.

Which signals score a passage independently of the query?

6 query-independent signals score the passage in the patent's example process, which states that "more scoring features, or fewer scoring features, can be used":

  1. Position: the higher the passage sits on the page, the higher the score.
  2. Language model: passages with complete sentences beat partial sentences, and a tri-gram model compares the passage with historical answer passages, "answer passages that have been served for all queries," because served answers share "a similar n-gram structure" as they "include explanatory and declarative statements." Structured content is exempt in some implementations: a table row can have a very low language model score yet be very informative, so it "is not subject to language model scoring."
  3. Section boundaries: a passage "will be penalized if it includes text that passes formatting boundaries, such as paragraphs and section breaks."
  4. Interrogative terms: a passage that contains a question scores lower than one with "only declarative statements." The example: "The moon is approximately 238,900 miles from the Earth" beats a passage that asks how far away the moon is.
  5. Discourse boundary terms: a passage that begins with "conversely," "however" or "on the other hand" receives a low score; one that contains such a term later scores higher, and one without it scores highest.
  6. Resource scores: the ranking score of the page for the query, its reputation score ("the trustworthiness and/or likelihood that that subject matter of the resource serves the query well") and the quality score of the site that hosts it. "Generally, the higher these scores are, the higher the answer score will be."

These components are summed, multiplied or combined in other ways into the query independent score.

How is the final answer score computed?

The final answer score combines the query dependent and the query independent scores, again by sum, product or another method. The patent leaves room for variants: "only the query dependent score may be used for the answer score," and the answer score "may be adjusted according to additional scoring processes." The passage with the best score is the one selected and shown with the results.

How does this patent relate to the heading context patent?

This patent scores the passage itself, while the context scoring patent adjusts that score with the headings above the passage. Steven D. Baker and Srinivasan Venkatachary signed both. The 2 patents agree on questions: a question inside the passage lowers its score, while a question just before the passage, ideally as its heading, raises it.

The resource scores include the site quality score, the kind of site-level signal described in the site quality score patent.

What does answer passage scoring change for your SEO?

Answer passage scoring means that the passage Google lifts is a short, declarative, self-contained answer that uses the vocabulary of the top results. 8 consequences follow:

  1. Rank in the top results first. In the described implementation, candidates come from the top N resources, and N can match the first page.
  2. Use the answer's vocabulary. Terms that most top results use carry the highest answer term weight. Cover them in the answer itself.
  3. Answer with declarative sentences. State the answer ("The moon is approximately 238,900 miles from the Earth"), and put the question in the heading, not in the answer.
  4. Keep each answer inside one section. Passages that cross paragraphs or sections are penalized, and units under different headings do not mix.
  5. Do not open an answer with "however." Discourse boundary terms at the start of a passage lower its score.
  6. Build complete lists with order terms. Step lists with "first," "second" or "steps" are taken whole, so keep them short enough to fit a passage, and put the key information in the first sentence of each list item.
  7. Name the entity the question expects. For a "who" question, a passage that names a person of the expected type avoids the entity type penalty. For these questions, a separate patent on entity answers counts the entities named in the top results.
  8. Use clean tables and key value pairs. Tables let the system pick the rows that answer superlative or attribute questions, and a short sentence starting with "These" can introduce the table.

Our semantic SEO work applies these 8 rules to each answer section, starting with the vocabulary the top results share for the question.

The patent describes what Google's system can do. It does not confirm that featured snippets use these exact signals today.

Quick quiz

Did you get it?

Test what you just read.

Question 1 of 4

Where do answer terms come from?

Related articles

We read the patents so you don't have to.

Every week: 3 data-backed SEO insights, 1 myth busted, 1 pattern to steal. No fluff. No guru advice.

Free forever. Unsubscribe anytime. We respect your inbox. Privacy.