Subscribe to Newsletter
PageRankPatentsLink building

Google's seed-based PageRank: why your distance from trusted sites shapes your SEO authority

By 8 min read

Seed-based PageRank is a ranking method that scores each page by its shortest link distance from a set of trusted seed pages. Patent US 9,165,040 B1, "Producing a ranking for pages using distances in a web-link graph," describes it. Google filed it on October 12, 2006, and the patent was granted on October 20, 2015.

What does patent US 9,165,040 describe?

Patent US 9,165,040 describes a PageRank variant that measures authority as closeness to trusted pages, not as the sum of all incoming links. Its sole inventor is Nissan Hajaj. It contains 34 claims and 3 drawing sheets. A continuation, US 9,953,049 B1, was filed on October 19, 2015, the day before the grant.

The patent starts from a known weakness of the original algorithm: "Some web pages (called 'spam pages') can be designed to use various techniques to obtain artificially inflated PageRanks, for example, by forming 'link farms' or creating 'loops.'" A known fix already existed: "One possible variation of PageRank that would reduce the effect of these techniques is to select a few 'trusted' pages (also referred to as the seed pages) and discovers other pages which are likely to be good by following the links from the trusted pages."

That fix had a cost. The patent notes that "this variation of PageRank requires solving the entire system for each seed separately," so that "as the number of seed pages increases, the complexity of computation increases linearly, thereby limiting the number of seeds that can be practically used." The invention replaces the per-seed PageRank computation with shortest distances, which the patent says can be computed together for all the pages and all the seeds.

What is a seed page?

A seed page is a high-quality page specially selected as a starting point for the ranking. The patent sets 3 criteria: seeds "need to be reliable, diverse to cover a wide range of fields of public interests, as well as well-connected with other pages (i.e., having a large number of outgoing links)." Its 2 examples are the Google Directory and The New York Times. Seeds with many useful outgoing links act as "hubs" on the web.

The patent wants many seeds: "it is desirable to use large number of seed pages to accommodate the different languages and a wide range of fields." A more diverse set of seeds "can shorten the paths from the seeds to a given page." It also states the limits: "because selecting the seeds involves a human manually identifying these high-quality pages, the total number of the seeds is typically limited. Moreover, having too many seeds can make the selected seeds vulnerable to manipulation."

Seeds are not all equal. Each seed can carry an optional weight w between 0 and 1 (1 by default), converted into an initial distance of −log(w): a seed with a lower weight starts the race with a handicap. A seed can also consist of more than one page, in which case the path starts from whichever of its pages is closest.

How does Google measure the distance between 2 pages?

Google measures the distance between 2 pages by adding the lengths of the links on the shortest path between them. The patent says the length of a link "can be a function of any set of properties of the link and the source of the link," which "can include, but are not limited to, the link's position, the link's font, and the source page's out-degree."

Its main model uses the number of outgoing links of the source page: L(q → p) = α + log(|q|out), where |q|out is the number of outgoing links of page q and α = −log(d), with d the damping factor of PageRank. The length "increases as the number of outgoing links from the source page q increases." The function can also include a weight of the link. The patent shows that other length functions work too, including a constant length of 1 per link, which amounts to counting clicks.

The distance from a seed to a page is the shortest sum of link lengths: D(p) = min (D(q) + L(q → p)) over all pages q that link to p, starting from the seed's initial distance. If no path exists, the distance is infinite. A link from a page with 5 outgoing links is shorter than a link from a page with 500 outgoing links.

How does distance become a ranking score?

Distance becomes a ranking score through a score proportional to e^(−D(p)): the shorter the distance, the higher the score. The exponential turns each link of the path into a multiplier of d/|q|out, which is the share of PageRank that a page passes through each of its links. The patent justifies keeping only the best path with an observation: "the incoming contributions for a page have a significantly skewed distribution such that the sum of the all the incoming contributions is dominated by one or very few terms." The score is therefore an approximation of seeded PageRank, built from the single strongest path.

The distance used is not the distance to the nearest seed but to the k-th nearest seed. The patent notes: "In practice, it is sufficient to choose k to be a small integer, for example, 3, 4, 5, or 6," and the claims require k to be greater than 1 and smaller than the number of seeds. This choice "facilitates suppressing unfairly high scores due to lack of proportionality at the vicinity of the seeds." A page that sits one link away from a single seed but far from every other seed gets the score of its k-th distance, not of its best one.

What happens to pages that seeds cannot reach?

Pages that seeds cannot reach receive no score at all. The patent is explicit: "a page that cannot be reached by any of the seed pages will not be ranked." Because the final distance is the k-th shortest one, the formula also gives an infinite distance to a page that fewer than k seeds can reach. A link farm can link to itself thousands of times; if no path leads to it from the seeds, its internal links produce no distance and no score.

The links that sit on these k shortest paths form a "reduced link-graph." It holds far fewer links than the full web graph, keeps the same shortest distances, and gives each page at most k incoming links. The rank flow to each page can be traced back to its k nearest seeds through it.

What else does the patent use the distances for?

The patent uses the distances to tune the seeds and to guide crawling, indexing and ranking:

  • Seed tuning. The ranking produces lists of the nearest seeds and the lengths of the shortest paths for every ranked page. The system can use them "to evaluate the quality and the contribution of the seeds, and then modify the list of seeds and/or the weights of the seeds."
  • Crawling. The web crawler "can prioritize the crawling process by using the page rank scores."
  • Search. Pages are indexed and ranked with this process, and the search engine uses the ranking information to identify highly ranked documents that satisfy a query. Query-specific authority belongs to another patent, which counts links from pages that already rank for the query.
  • Beyond the web. The technique applies to "any hyperlinked database," including "hyperlinked documents of an enterprise."

How does seed-based PageRank compare with the original PageRank?

Seed-based PageRank compares with the original PageRank on 4 points:

PointOriginal PageRankSeed-based PageRank
Starting pointEvery page of the webA set of trusted seed pages
What countsEvery path, summedThe shortest path from the k-th nearest seed
Link farms and loopsCan inflate scoresEarn nothing without a path from the seeds
Unreachable pagesReceive a base scoreReceive no score

The 2 models share one rule: a page with many outgoing links passes less through each of them. The reasonable surfer model refines that rule further by weighting each link by its chance of being clicked.

What does seed-based PageRank change for your SEO?

Seed-based PageRank means that who links to you matters less than how close those links bring you to trusted sites. 5 consequences follow:

  1. Earn links from pages close to the seeds. The patent's example seeds are a major directory and a major newspaper, and it assumes seeds are "closer" to other high-quality pages. A link from a page a few links away from such sources shortens your distance more than dozens of links from pages no seed reaches.
  2. Prefer links from focused pages. In the main model, link length grows with the log of outgoing links, so a link from a page with a few outgoing links is shorter than a link from a directory page with hundreds. The patent also lists the link's position and font as possible length factors.
  3. Be close to several trusted sites, not one. The score uses the k-th nearest seed, with k a small integer (3 to 6 in the patent's examples): proximity to a single authority is not enough.
  4. Keep every page reachable. A page that no path reaches is unranked, so a page with no link pointing to it gets nothing. Each extra click from your homepage adds length to the path.
  5. Ignore link farms and link exchanges. Loops and closed networks of sites earn nothing when no short path leads to them from the seeds. Google can detect these structures directly: link farms and link rings leave a mathematical signature in PageRank.

Trust as a ranking signal goes beyond link distance: the trust patent describes how Google can rank results by the users and sources that vouch for them.

The patent describes what Google's system can do. It does not confirm which pages Google uses as seeds today, or that this method replaced the original PageRank.

Quick quiz

Did you get it?

Test what you just read.

Question 1 of 4

What determines the score of a page in seed-based PageRank?

Related articles

We read the patents so you don't have to.

Every week: 3 data-backed SEO insights, 1 myth busted, 1 pattern to steal. No fluff. No guru advice.

Free forever. Unsubscribe anytime. We respect your inbox. Privacy.