Subscribe to Newsletter
PageRankInternal linkingArchitecture

How SEO authority actually flows: internal linking and Google's PageRank patent

By 9 min read

PageRank is a score that measures the importance of a page from the links pointing to it, weighted by the importance of the pages that cast those links. Patent US 6,285,999 B1, "Method for node ranking in a linked database," describes it. Lawrence Page filed it on January 9, 1998, claiming priority from a provisional application filed on January 10, 1997, and the patent was granted on September 4, 2001.

What does the PageRank patent describe?

The PageRank patent describes a method that assigns an importance rank to every node of a linked database: the web, but also any database of documents that cite each other. Its sole inventor is Lawrence Page, and the assignee is The Board of Trustees of the Leland Stanford Junior University. The patent has 29 claims and 3 drawing sheets. It has since expired.

The abstract states the core idea: "The rank assigned to a document is calculated from the ranks of documents citing it." The patent contrasts it with simple citation counting, which gives every citation the same value: "A citation from an important document is more important than a citation from a relatively unimportant document."

How is a page's rank calculated?

A page's rank is calculated from the ranks of the pages that link to it, each divided by the number of forward links on the linking page. The patent's formula:

r(A) = α/N + (1 − α) · (r(B1)/|B1| + … + r(Bn)/|Bn|)

  • B1 … Bn are the backlink pages of A, and r(B1) … r(Bn) are their ranks.
  • |B1| … |Bn| are their numbers of forward links.
  • α is a constant between 0 and 1: the probability that the surfer "will jump randomly to any web page instead of following a forward link."
  • N is the total number of pages in the web.

SEO literature usually writes the same formula with a damping factor d = 1 − α: PR(A) = (1 − d)/N + d · Σ PR(i)/L(i). The patent itself calls (1 − α) the damping factor, which "limits the extent to which a document's rank can be inherited by children documents." It gives α as "typically around 15%" in one passage and "~0.1" in its illustrative example.

The ranks form a probability distribution: they sum to 1 over all pages. "The rank of a page can be interpreted as the probability that a surfer will be at the page after following a large number of forward links."

One strong link can beat many weak ones because each backlink passes on the rank of the linking page, not a fixed vote. The patent says it directly: "it is possible, therefore, for a document with only one backlink (from a very highly ranked page) to have a higher rank than another document with many backlinks (from very low ranked pages)." A mass of weak backlinks has a cost beyond its low value: the math of PageRank lets Google detect link spam such as link farms.

It adds a condition. A citation from a highly ranked backlink counts more than a citation from a lowly ranked one "provided both citations come from backlink documents that have an equal number of forward links." The number of links on the linking page matters as much as its rank: a page that links to 2 pages gives each one half of its share, a page with 150 links gives each one 1/150.

What does the three-page example show?

The patent's example (FIG. 2) uses 3 documents: A links to B and C, B links to C, and C links to A.

  • With α = 0 (no random jump), the ranks are r(A) = 0.4, r(B) = 0.2 and r(C) = 0.4. A receives all of C's rank, because A is C's only forward link. B receives only half of A's rank, because A has 2 forward links.
  • With α = 0.5, the equations become r(A) = 1/6 + r(C)/2, r(B) = 1/6 + r(A)/4 and r(C) = 1/6 + r(A)/4 + r(B)/2. The solution is r(A) = 14/39, r(B) = 10/39 and r(C) = 15/39.

With the random jump, C, the only page with 2 backlinks, moves ahead of A. The value of α changes not only the scores but the order.

How is PageRank computed across millions of pages?

PageRank is computed by iteration. As the initial state, the system "may simply set all the ranks equal to 1/N," then uses the formula to compute new ranks from the existing ones. "In the case of millions of documents, sufficient convergence typically takes on the order of 100 iterations." The patent adds that "even approximate rank values, using two or more iterations, can provide very valuable, or even superior, information."

The patent describes the same computation as a random surfer model (FIG. 3). A transition matrix gives the probability of moving from page i to page j, and the steady state is the dominant eigenvector of that matrix. "The iteration circulates the probability through the linked nodes like energy flows through a circuit and accumulates in important places."

2 details matter for site owners:

  • Pages without forward links. These childless pages "bleed off energy" and complicate the computation. The patent proposes removing them during the iterations and adding them back once the iteration is complete, then running as many iterations again so that they all receive a value.
  • Loops. Damping is "important when many iterations are used to calculate the rank so that there is no artificial concentration of rank importance within loops of the web."

Which variants does the patent describe?

The patent describes several adaptations of the base method, and several of them concern links between pages of the same site:

  1. Local links ignored or discounted. "A modification to avoid drawing unwarranted attention to pages with artificially inflated relevance is to ignore local links between documents and only consider links between separate domains." A simpler approach "is to weight links from pages contained on the same web server less than links from other servers."
  2. Diverse sources. "Rank can be increased for documents whose backlinks are maintained by different institutions and authors in various geographic locations."
  3. Important locations. Rank can be increased "if links come from unusually important web locations such as the root page of a domain."
  4. Visible links. "Highly visible links that are near the top of a document can be given more weight," and so can links in large fonts or emphasized in other ways.
  5. Fresh links. Higher value can go to links "coming from pages that have been modified recently since such information is less likely to be obsolete."
  6. Targeted random jumps. The random jump can lead only to a few high-importance nodes, which "can be very effective in preventing deceptively tagged documents from artificially inflated relevance." Google's trusted seeds patent builds on the idea of starting from trusted pages.
  7. Personalization. A user's home page or bookmarks can receive a large initial importance or a high probability that a random jump returns to them.

The visible-links variant foreshadows the reasonable surfer patent, which weights each link by its probability of being clicked.

The claims cover these factors. Claims 2 to 7 weight linking documents by their number of links, the probability that they will be accessed, their "URL, host, domain, author, institution, or last update time," the importance, visibility or textual emphasis of their links, or a user's preferences. Claim 10 covers ranking documents by an automated random traversal, counting how many times each one is traversed.

How does the search engine use the rank?

The search engine uses the rank as one factor combined with text matching. A crawler builds an index of the content and a graph of the links. The engine finds the documents that match the query, in the full text, in the titles or in "the anchor text associated with backlinks to the page," then sorts them "with high ranking documents first." The patent specifies: "The ranking in this case is a function which combines all of the above factors such as the objective ranking and textual matching."

In the patent, the rank orders documents that already match the query. It does not replace the match.

What does PageRank change for your SEO?

PageRank means that a link passes a share of the linking page's rank, divided by the number of links on that page. 6 consequences follow:

  1. Get links from important pages. One link from a highly ranked page can outweigh many links from weak pages.
  2. Limit dilution on the pages that matter. Every forward link takes a share. A page with 150 links passes 1/150 of its share through each one. Trim boilerplate links on pages meant to push rank to your key pages.
  3. Link from your strongest pages. Rank flows from the pages that already have it. Your highest-ranked pages, often the homepage and the most-linked articles, are your best internal-link sources. Point them at the pages you want to rank.
  4. Build clusters, not chains. Damping limits how much rank a page passes on at each step, so rank fades along a long chain of pages. Group related pages and interlink them around a hub page: under the formula, the hub collects rank from the pages of the cluster and passes it to the pages it links to.
  5. Do not count on internal links alone. The patent proposes ignoring local links or weighting links from the same server less. Internal links organize how rank circulates inside a site; links from other domains bring it in.
  6. Avoid dead ends. Pages without forward links are a special case in the computation. Give every page links to related content.

Our internal linking audit measures how many clicks separate each key page from the homepage and how many links share the rank of the pages that point to it.

The patent describes the method as filed in 1998. It does not say how much weight Google gives PageRank today, or which of its variants Google applies.

Quick quiz

Did you get it?

Test what you just read.

Question 1 of 3

What does PageRank model?

Related articles

We read the patents so you don't have to.

Every week: 3 data-backed SEO insights, 1 myth busted, 1 pattern to steal. No fluff. No guru advice.

Free forever. Unsubscribe anytime. We respect your inbox. Privacy.