Subscribe to Newsletter
Anchor textPatentsCrawling

Anchor text indexing for SEO: how Google indexes a page with the words other pages use to link to it

By 7 min read
Anchor text indexing: Google's crawler patent and SEO, patent US 7,308,643 B1Anchor text

Anchor text indexing is the practice of indexing a page with the anchor text, and the text around it, that other pages use when they link to it. Patent US 7,308,643 B1, "Anchor tag indexing in a web crawler system," describes how Google's crawler collects that text and attaches it to target pages at web scale. Google filed it on July 3, 2003, and the patent was granted on December 11, 2007.

What does patent US 7,308,643 describe?

Patent US 7,308,643 describes a crawling and indexing system that records every link with its text, sorts the links by target page, and indexes each target with the text pointing to it. Its 5 inventors are Huican Zhu, Jeffrey Dean, Sanjay Ghemawat, Bwolen Po-Jen Yang and Anurag Acharya. Jeffrey Dean and Sanjay Ghemawat later co-authored Google's MapReduce and Bigtable papers. The patent, assigned to Google Inc. at grant, holds 27 claims and 13 drawing sheets, and its term was extended by 587 days.

The abstract states the core: "A link log, including one or more pairings of source documents and target documents is accessed. A sorted anchor map, containing one or more target document to source document pairings, is generated," with pairings "ordered based on target document identifiers."

The crawler records links in link logs. For each crawled page, a record holds "the fingerprints of all the links (URLs) that are found" in it, "as well as the text that surrounds the link." The patent's example: the text "to see a picture of Mount Everest click here," where the link points to an image of Mount Everest.

The patent describes a crawl organized in layers. URLs are placed in a base layer, a daily layer or a real-time layer "based on the historical (or expected) frequency of change of the content of the web pages at the URLs and a measure of URL importance," such as page rank. In one embodiment, a crawl cycle, or epoch, lasts one day, and real-time, daily and base indexers make documents searchable.

Link maps and anchor maps are 2 views of the same link logs, keyed differently:

  1. Link maps are keyed by the source page and list its targets, without text. Page rankers use them "to adjust the page rank of URLs."
  2. Anchor maps are keyed by the target page and hold "the text that corresponds to the URL." Indexers use them "to facilitate the indexing of 'anchor text' as well as to facilitate the indexing of URLs that do not contain words."

Because anchor records are sorted by target, "binary search techniques can be used to quickly locate the record corresponding to the particular target." Both are built as layered sets of sorted maps, merged when a merge condition is met (a time schedule, too many maps, too much data, or idle time). For link maps, the patent prefers merging maps of similar size, "within a factor of 2 of each other."

The 2 maps are read differently. Page rankers keep only the most recent record for a source page. Indexers "simply take all the information available," so they read every anchor map that contains the target page.

When a link disappears, its text stops being associated with the target page. The global state manager writes a "delete link entry" when it "determines that a link no longer exists," for example when an older record of a source page lists a target that a newer record no longer lists. When a whole document has been removed, it writes a "delete node entry." During a merge, the most recent map wins: in the patent's example, because the map holding the delete entry is more recent, the source page is not included in the merged anchor record. For each source and target pair, the merged map keeps "the most recent annotation."

Why index a page with text from other pages?

Indexing a page with text from other pages describes the page even when its own content cannot. The patent gives 3 cases:

  1. The page is unavailable: if the target "is unavailable for retrieval" when the crawl is performed, because its server is down or asks for a password, "this anchor text provides textual information that can be searched by keyword."
  2. The page has no text: the target "may be an image file, a video file, or an audio file," with no text available. The Mount Everest image has no words, but the link text describes it.
  3. Others describe it better: "document 1002 contains more accurate information about document 1012-1 than the textual contents of document 1012-1 itself." The patent's example: an authoritative page states that "the server that hosts web page 1012-1 is frequently unavailable," while the page itself "may contain no text indicating that it is unavailable."

The patent's simplest example: if the anchor text "this is an interesting website about cats" is indexed as part of the target page, "a user who submits a query containing the term 'cat' may receive a list of documents including" that page. The summary adds a further advantage: "the ability to index a web page before the web page has been crawled." Another Google patent scores anchor text against the related phrases found on the linking page and on the target page.

What information is stored with anchor text?

Anchor text is stored as annotations with attributes. The patent lists formatting attributes of the surrounding text, such as text emphasized with <EM>, a citation with <CITE>, a variable name with <VAR>, strongly emphasized text with <STRONG> and source code with <CODE>. Other attributes include "text position, number of characters in the text passage, number of words in the text passage."

The stored text is not limited to the anchor itself. A text passage can be taken "from text within a predetermined distance of an anchor tag." That distance can depend on "a number of characters in the HTML code of the source document, the placement of other anchor tags in the source document," or other "anchor text identification criteria." At query time, the search engine can search the page's own content and also "any annotations associated with a document for one or more of the query terms."

How does the patent handle duplicate pages?

The patent pools the anchor text of duplicate pages. Before indexing, a duplicate server compares duplicates by page rank and identifies the "canonical" page. A non-canonical page is not forwarded for indexing. In some embodiments, the entry for a page lists the URL fingerprints of its duplicates, limited to K entries, with K "preferably having a value between 2 and 10." The indexers then read "the anchor text for the links pointing to each of the identified duplicate pages and index that anchor text as part of the process of indexing the page." The patent notes this is useful when links to a non-canonical page carry "anchor text in a different language than the anchor text of the links to the canonical page."

The words around a link carry meaning in other Google patents: the link context patent fingerprints them to detect spam, and the reasonable surfer patent uses them to estimate how likely a link is to be clicked.

What does anchor text indexing change for your SEO?

Anchor text indexing means that the words others use to link to you become part of how Google describes your page. 6 consequences follow:

  1. Earn descriptive links. Anchor text such as "an interesting website about cats" can make the target findable for "cat": the description others give you is indexed with your page.
  2. Mind the sentence around the link. The crawler stores "the text that surrounds the link," not only the anchor, within a predetermined distance: a link inside a descriptive sentence carries more meaning than a bare "click here."
  3. Describe your images and media with links. Images, video and audio have no text of their own: internal links with descriptive text tell the index what they show.
  4. Know that others can describe you. An authoritative page can attach facts to your page, positive or negative, that your own content does not contain.
  5. Link internally with intent. Your internal anchors are anchor text too: use them to describe each target page accurately.
  6. Keep canonical signals clean. Links pointing to duplicates of your page can feed its anchor text, but only the canonical version is indexed: make sure the version you want is the one chosen.

The patent describes what Google's system can do. It does not confirm how Google weighs anchor text in its rankings today.

Quick quiz

Did you get it?

Test what you just read.

Question 1 of 4

What does the crawler store for each link in the link logs?

Related articles

We read the patents so you don't have to.

Every week: 3 data-backed SEO insights, 1 myth busted, 1 pattern to steal. No fluff. No guru advice.

Free forever. Unsubscribe anytime. We respect your inbox. Privacy.