A link context is a fingerprint of the rare words that surround a link in the page that contains it. Patent US 8,577,893 B1, "Ranking based on reference contexts," ranks a page by the contexts of all the links pointing to it, to reward varied, natural citations and discount mass-produced ones. Google filed it on March 15, 2004, and the patent was granted on November 5, 2013.
What does patent US 8,577,893 describe?
Patent US 8,577,893 describes a ranking factor based on the context of the links that point to a document. Its 2 inventors are Anna Patterson, the author of the phrase-based indexing patent, and Paul Haahr. The patent holds 17 claims and 9 drawing sheets, and its term was extended by 2,935 days under 35 U.S.C. 154(b).
The abstract sums it up: the system "identifies a rare word (or words)" in the part of a page around a link, "creates a context identifier based on the rare word(s), and ranks the second document based on the context identifier." The patent defines a link broadly as "any reference to or from a document."
Which manipulations does the patent target?
The patent targets 4 techniques that inflate rankings artificially:
- Link-based spamming: obtaining many links to a page, through link farms "heavily cross-linked to each other" or by paying owners of high-ranking pages to add a link.
- Anchor text spamming: many pages linking to a document "using the same anchor text with which the document is to be associated."
- Bombing: the patent cites the bomb in which many pages used the anchor text "miserable failure" to link to President Bush's biography, which then ranked first for that query.
- Standard frames: links repeated on every page of a site, such as "products," "jobs" or "investor" links, which "may artificially inflate the ranks" of their targets "especially when the web site includes a large number of documents."
The patent also lists the motives: financial ones, since "many document owners will pay to have their document highly ranked," but also "disrupting the search engine, humiliating the search engine company, or no reason at all." Its answer to all 4 techniques is the same: use the context of the links as a ranking factor, so that artificially inflated rankings "may be reduced."
How does Google build a link context?
Google builds a link context in 4 steps for each link it finds while parsing a document during crawling or indexing. The parsed document is a different document from the one being ranked:
- Read 2 windows of text: one to the left of the link and one to the right. The default window is 5 words; the patent mentions more or fewer, such as 15.
- Find the rarest word in each window, with "an inverse document frequency (IDF) weighting technique or a conventional linguistic modeling technique." One technique hashes every word of the corpus into a table whose counts show which words occur less often. In one embodiment, the rarest words are "real" words, for example a word that occurs "at least fifty times on many different documents," which excludes random blocks of text with symbols or numbers. Another implementation picks a rare phrase (a combination of words) in each window instead of a single word.
- Hash the 2 rare words into a context identifier, a fingerprint of the link's context. Each word "may be hashed individually or after combining them."
- Count each context across all the links pointing to the document: the context count is the number of times those rare words occurred in association with links to the document.
The patent's example is a page about Saturn with the sentence "Perhaps the most beautiful of all the planets, Saturn is surrounded by an elegant and intriguing ring system." The anchor text is "Saturn," linking to www.planetsaturn.com. The left window is "beautiful of all the planets," the right window is "is surrounded by the elegant," the rarest words are "planets" and "elegant," and their hash gives the context 112. Queries get a comparable reading of surrounding words: Google checks synonyms in long queries against terms far from the word being replaced.
How do contexts change the ranking of a page?
Contexts change the ranking through 3 measures computed from the list of contexts and their counts:
- The number of different contexts: the number of entries in the list is used "to determine a ranking score for the document" (claim 5 speaks of "a total number of the context identifiers"). The patent does not state the direction of this factor, but its anti-spam examples imply that many distinct contexts weigh more than one context repeated thousands of times.
- The distribution of context counts: a context with a disproportionate count is suspicious. In the patent's example, contexts with counts of 10,000, 10, 4 and 1: the first one is discounted "as suspicious (e.g., possibly machine generated)," and the page is ranked as if it had 3 contexts.
- The distribution history: a sudden change is suspicious. A page with 2 contexts counting 20 each gains a third context counting 18,000 in the next period: the jump from a total of 40 to 18,040 marks the document "as suspicious."
In the Saturn example, the list holds contexts 23, 46, 112 and 156 with counts of 30,000, 15, 8 and 3. Context 23 is judged suspicious "due to its disparate context count compared to the other contexts," then "eliminated from, or its contribution reduced in, the ranking process." The examples flag a single context, but the patent adds that "multiple contexts may be deemed suspicious." The ranking can also be pre-calculated for each document and simply looked up when a query arrives.
The context factor combines with "the number of links to the document, the importance of the documents linking to the document, the freshness of the documents linking to the document, and/or other known ranking factors."
Does the context replace anchor text?
No: the context complements anchor text. The patent states that other data, "text or non-text," can define the context, such as "data surrounding the link, data to the left of the link or to the right of the link, or anchor text associated with the link." In the patent's logic, a bomb can control the anchor text, but thousands of copies of the same sentence produce one context with a huge count, which is precisely what the distribution measure discounts.
The surrounding words of a link matter in other Google patents: the reasonable surfer patent weighs a link by its chance of being clicked, using features that include "the context of a few words before and/or after the link."
What do the 17 claims protect?
The claims protect the rare-word version of the mechanism. Claim 1 covers identifying a link, analyzing the text to its left and to its right, picking a rare word in each portion "based on a frequency of occurrence" in a set of documents, creating a context identifier "based only on the first and second rare words," and ranking the target "within a list of search results." The dependent claims add:
- Hashing the 2 rare words into the identifier (claims 3 and 13).
- The total number of context identifiers as a ranking basis (claim 5).
- The distribution of context counts, with the impact of one identifier reduced (claims 7 and 8).
- The history of that distribution (claim 9).
- A ranking score used as "one of a plurality of factors" (claims 10, 14 and 17).
Claim 15 extends the system to a "rare word or rare phrase" on each side of the reference. Claim 17 covers the whole set of different contexts gathered from many linking documents.
What does the link context patent change for your SEO?
The link context patent means that a link profile is worth more when links come from varied, natural sentences than when thousands of links share the same text. 5 consequences follow:
- Earn links in varied editorial contexts. The number of distinct contexts feeds the ranking score: citations in different articles, written by different people, create different contexts.
- Avoid templated link placements. Guest posts, widgets or press releases that repeat the same sentence around your link create one context with a disproportionate count.
- Watch sitewide links. Footer and navigation links repeated on every page are the patent's "standard frames," and the same surrounding words on every page can produce a single repeated context.
- Grow links at a natural pace. A new context that appears 18,000 times in one period, on a page that had a total of 40, is the patent's example of a document flagged as suspicious.
- Let the surrounding text vary naturally. The rare words around a link define its context. The patent counts these contexts and their distribution; it does not score whether the words match your topic, so variety matters more than a scripted sentence.
The patent describes what Google's system can do. It does not confirm that Google uses link context fingerprints in its rankings today.