Subscribe to Newsletter
Phrase-based indexingPatentsSemantic SEO

What phrase-based indexing means for SEO: how Google learned that topics are made of related phrases

By 13 min read

Phrase-based indexing is a retrieval method that indexes and ranks documents by meaningful phrases and by the related phrases that predict them, instead of by isolated words. Patent US 7,536,408 B2, "Phrase-based indexing in an information retrieval system," describes it. Google filed it on July 26, 2004, and the patent was granted on May 19, 2009.

What does patent US 7,536,408 describe?

Patent US 7,536,408 describes a search system that uses phrases to index, retrieve, organize and describe documents. Its sole inventor is Anna Lynn Patterson, then an engineer at Google. The abstract states the core idea: "Phrases are identified that predict the presence of other phrases in documents." Documents are then indexed according to the phrases they contain. The patent has 20 claims and 8 drawing sheets.

The patent covers 6 uses of phrases:

  1. Indexing: each document is listed under the phrases it contains, with a record of the related phrases also present.
  2. Retrieval: the phrases of a query, and their phrase extensions, select the candidate documents.
  3. Ranking: related phrases decide how well a document covers the query's topic.
  4. Clustering: documents are grouped by topic in the results.
  5. Description: the sentences richest in query phrases and related phrases form the document description.
  6. Duplicate elimination: documents whose key sentences match are treated as duplicates.

The patent also describes an optional personalization of results, based on a user model made of phrases. Its claims protect the indexing step: storing, in the posting list of each phrase, the document identifier and an indication of each related phrase also present in the document. Claims 17 to 20 add a link score, computed from related phrases, when the phrase is the anchor text of a link.

The patent belongs to a large family. Applications were filed with 9 other patent offices, including the European Patent Office, China, Japan, Korea and Canada, and 3 US continuations extend it. Google filed it on the same day as 6 related US applications on phrase identification, phrase-based searching, personalization, taxonomy generation, document descriptions and duplicate detection.

What is a good phrase?

A good phrase is a sequence of words that is used often enough, or prominently enough, and that predicts the presence of other phrases. The system crawls the collection in partitions of preferably about 1,000,000 documents. It reads every document with a phrase window "preferably 4 or 5 terms (words)" long, stop words such as "a" and "the" included, and every sequence that starts at the current word is a candidate phrase. In one embodiment, a candidate moves from the possible phrase list to the good phrase list when it meets one of 2 conditions:

  • it appears in more than 10 documents and more than 20 times in total, or
  • it has more than 5 "interesting" instances, that is, instances set apart by formatting or grammar.

These thresholds scale with the partition: with 2,000,000 documents per partition, "the thresholds are approximately doubled." A phrase is a bad phrase when it appears in fewer than 2 documents and has no interesting instance. The good phrase list naturally includes single words, so the system indexes single words and multiword phrases alike.

Frequency is not enough. Good phrases "are predictive of other good phrases, and are not merely sequences of words that appear in the lexicon." "President of the United States" predicts phrases such as "George Bush" and "Bill Clinton." The patent gives idioms as counterexamples: "fell down the stairs," "top of the morning" and "out of the blue" appear with many unrelated phrases and predict nothing.

The system then prunes the list in 2 ways. A phrase that predicts no other good phrase is removed. An incomplete phrase, one "that only predicts its phrase extensions," moves to an incomplete phrase list: "President of the United" only predicts "President of the United States," so it leaves the good phrase list. At query time, this incomplete phrase list lets the system suggest or search the most likely extension. In a typical embodiment, the good phrase list holds "about 6.5 × 10^5 phrases," or 650,000. Because the crawl repeats as new documents arrive, the system detects new phrases as they enter the lexicon.

How does Google measure that one phrase predicts another?

Google measures prediction with information gain: the ratio between how often 2 phrases actually co-occur and how often chance alone would make them co-occur. The expected rate is the product of the shares of documents that contain each phrase. The actual rate is the number of co-occurrences divided by the total number of documents. Co-occurrences are counted in a secondary window around each phrase, 30 words in the patent's example.

The patent uses 2 thresholds:

  • 1.5 for prediction: one phrase predicts another when their information gain exceeds a threshold. "In one embodiment, the information gain threshold is 1.5, but is preferably between 1.1 and 1.7." Every good phrase must predict at least one other good phrase.
  • 100 for a related phrase: 2 phrases are related when their information gain exceeds 100, a co-occurrence "well beyond the statistically expected rates." In the patent's example, given "Monica Lewinsky" in a document, "Bill Clinton" is 100 times more likely to appear in it than in a randomly selected document.

The system also tracks "interesting" instances: a phrase "in boldface, or underline, or as anchor text in a hyperlink, or in quotation marks" stands out from its surroundings and is counted separately. For each pair of phrases, the system counts how often either one, or both, appear as distinguished text. The count of both is designed to avoid treating as predictive a phrase that repeats in sidebars, footers or headers, such as a copyright notice.

Related phrases are phrases that people use together to discuss the same topic. The patent's example is "President of the United States" and "White House." Each good phrase receives an ordered list of related phrases, from the highest information gain to the lowest.

A cluster is "a set of related phrases in which each phrase has high information gain with respect to at least one other phrase." One phrase can belong to several clusters, and a relationship can be mutual or one-directional. The patent's example finds 4 clusters among "Bill Clinton," "President," "Monica Lewinsky" and "purse designer": for instance, "Bill Clinton," "President" and "Monica Lewinsky" form one cluster, while "Monica Lewinsky" and "purse designer" form another. Each cluster receives a cluster number and takes the name of its related phrase with the highest information gain.

How does Google determine the topics of a document?

Google determines the topics of a document from the related phrases and secondary related phrases it contains. For each phrase of a document, the posting list stores 2 bits per related phrase, ordered by decreasing information gain. The first bit records whether the related phrase is present in the document. The second bit records whether one of that related phrase's own related phrases, a secondary related phrase, is present too.

A pair (1,1) marks a primary topic: according to the patent, the author "used several related phrases" together in drafting the document. A pair (1,0) marks a secondary, less significant topic.

How does Google rank documents with phrases?

Google ranks documents by the related phrases of the query that they contain, weighted by their information gain. The most significant bits of the vector belong to the strongest related phrases, so the value of the vector itself can serve as a score: "documents that contain high order related phrases of a query phrase are more likely to be topically related to the query than those that have low ordered related phrases." Such documents can rank well "even if the documents do not contain a high frequency of the input query terms." The patent describes 3 scores:

  1. Body hit score: the numerical value of the highest valued related phrase bit vector of the document for the query phrases. In a second embodiment, related phrases earn points instead: N points for the strongest related phrase, N-1 for the next, down to 1 point for the last.
  2. Anchor hit score: for each page that links to the document with a query phrase as anchor text, the related phrase bit vector of the anchor phrase in the linking page is multiplied by the one of the document. Linking pages that are themselves about the query phrase raise the score.
  3. Combined score: a linear combination of the 2, for example "Score=0.30*(body hit score)+0.70*(anchor hit score)," with weights that "can be adjusted as desired."

The system can then filter documents that cover too many topics, which "is particularly the case for longer documents." It removes documents that contain more than a threshold number of clusters. In the patent's example, it can "remove any documents that contain more than two clusters," because users often "prefer documents that are strongly on point with respect to a single topic." The threshold can be predetermined or set by the user.

How does Google use anchor text in phrase-based indexing?

Google scores each link by the related phrases of its anchor text found on the linking page and on the target page. When the anchor text of a link is a good phrase, the indexing system computes an outlink score from the related phrases present in the linking page, and an inlink score from the related phrases present in the target page. When the target page does not contain the anchor phrase itself, the system still checks which of its related phrases the target contains.

The patent's example is a link with the anchor "Australian Shepherd." The linking page contains only 1 of the 5 related phrases, "Aussie," so it is "only weakly about Australian Shepherds." The target page contains "blue merle," "red merle" and "tricolor." According to the patent, importing the target's related phrases counters link "bombing," where many pages with the same anchor text point to a page that "has little or nothing to do with the anchor text." A separate patent on anchor text indexing goes further and indexes the target page with the words of the anchor.

How does Google group search results by topic?

Google groups results by the related phrases of the query that appear in the most documents. The patent notes that "most users however, do not review beyond the first 30 or 40 documents," so relevant documents on other subtopics stay unseen. For each related phrase of the query phrases, the system counts how many result documents contain it. The most frequent related phrase names the first cluster, and so on for the top 3 to 5 clusters.

In the patent's example, for the query "blue merle agility training," "weave poles" appears in 75 of 100 documents and "teeter" in 60, so they name the first 2 clusters. The system can show a fixed number of documents per cluster, such as 10, or a number proportional to each cluster's size.

How does Google write snippets from phrases?

Google writes snippets by ranking the sentences of a document by the number of query phrases, related phrases and phrase extensions they contain. The first sort key is the count of query phrases, the second the count of related query phrases, the third the count of phrase extensions. The system selects the top sentences, "e.g., five sentences," to form the document description. A sentence that contains the query phrase and several of its related phrases is therefore the most likely to be selected. In the personalized version, the related phrases found in the user model come first.

How does Google detect duplicate documents with phrases?

Google detects duplicates by comparing the sentences that carry the most related phrases in each document. For each document, the system ranks sentences by the frequency of its related phrases, keeps the top N (e.g., 5 to 10), concatenates them and hashes the result. When 2 documents produce the same hash, they are duplicates. The system keeps the document with the higher page rank, or another query independent measure, and can remove the other one from the index "so that it will not appear in future search results for any query." The same check runs during crawling. The patent's example is a news agency article replicated on a dozen or more newspaper websites. Pages that differ by a few words fall under another method, in which Google fingerprints near-duplicate content to keep a single version.

Can phrase-based indexing detect spam?

Yes: a companion patent uses the same phrase statistics to detect spam. Patent US 7,603,345 B2, "Detecting spam documents in a phrase based information retrieval system," also by Anna Lynn Patterson, filed on June 28, 2006 and granted on October 13, 2009, identifies a spam document "based on the number of related phrases included in a document." It flags documents with an excessive number of related phrases, a statistically significant deviation from the expected number. A page stuffed with every related phrase of a topic is exactly the pattern it targets. Language models give Google a second way to catch keyword stuffing, described in the patent on gibberish content.

What does phrase-based indexing change for your SEO?

Phrase-based indexing means that a page ranks for a topic when it uses the phrases that define that topic, in natural proportions. 6 consequences follow:

  1. Cover the strongest related phrases. The related phrases with the highest information gain carry the most weight. List the phrases that expert pages on your topic use together, and cover them.
  2. Keep one topic per page. A document spread over too many clusters can be filtered out, more than 2 in the patent's example. Split a page that covers several topics into several pages.
  3. Get links from on-topic pages, with topic phrases as anchors. The anchor hit score combines the related phrases of the anchor phrase on the linking page and on the target page. A topical anchor on an off-topic page, or pointing to a page without the related phrases, earns little.
  4. Write sentences that can become descriptions. Combine the main phrase and its strongest related phrases in the same sentence.
  5. Keep your key sentences original. Pages whose top sentences match another document are treated as duplicates, and only the one with the higher page rank or other query independent measure is kept.
  6. Never stuff related phrases. The spam patent flags pages with an excessive number of related phrases. Natural coverage beats exhaustive lists.

Our semantic SEO work starts from the related phrases: we list the phrases that expert pages on a topic share, then split pages that spread over more than one cluster.

The same logic underlies entity-based SEO and the Knowledge Graph: topics are networks of connected concepts, and Google reads the network. The language of a whole site carries a quality signal as well, as the site quality patent shows with its n-gram model.

The patent describes what Google's system can do. It does not confirm that Google uses these exact thresholds today.

Quick quiz

Did you get it?

Test what you just read.

Question 1 of 4

Besides frequency, what makes a phrase a good phrase?

Related articles

We read the patents so you don't have to.

Every week: 3 data-backed SEO insights, 1 myth busted, 1 pattern to steal. No fluff. No guru advice.

Free forever. Unsubscribe anytime. We respect your inbox. Privacy.