The reasonable surfer is a link model in which each link passes authority according to the probability that a real user clicks it. Patent US 7,716,225 B1, "Ranking documents based on user behavior and/or feature data," describes it. Google filed it on June 17, 2004, and the patent was granted on May 11, 2010.
What does the reasonable surfer patent describe?
The reasonable surfer patent describes a version of PageRank in which links carry different weights. Its 3 inventors are Jeffrey A. Dean, Corin Anderson and Alexis Battle, and it was assigned to Google Inc. with 19 claims and 7 drawing sheets. The system builds a model from the features of links and from the links users actually follow, then uses the model to weigh each link in the rank computation.
The family continues with 3 continuations: US 8,117,209 B1, US 9,305,099 B1 and US 10,152,520 B1. The patent itself names the model: "This reasonable surfer model reflects the fact that not all of the links associated with a document are equally likely to be followed."
How does the reasonable surfer differ from the random surfer?
The reasonable surfer differs from the random surfer because it does not follow every link with the same probability. The original PageRank, Google's link authority patent, models a random surfer who picks any link on a page with equal chance, so a page with 100 links splits the authority it passes into 100 equal shares.
In the patent, "when a surfer accesses a document with a set of links, the surfer will follow some of the links with higher probability than others." It gives 3 examples of links unlikely to be followed: "'Terms of Service' links, banner advertisements, and links unrelated to the document." The share a link passes therefore depends on its chance of being clicked, not only on the number of links on the page.
The rank formula (Eqn. 1) keeps the PageRank structure and adds a weight w to each link: r(A) = a/N + (1 − a) · (w1 · r(B1)/|B1| + … + wn · r(Bn)/|Bn|). B1 to Bn are the documents linking to A, |B| is the number of forward links on each of them, a is a constant between 0 and 1, and N is the total number of documents. "A link's weight may reflect the probability that the link will be selected." The rank of a document is then "the probability that a reasonable surfer will access the document after following a large number of forward links."
What does the patent's worked example show?
The patent's example (FIG. 7) shows two links from the same page passing different amounts of rank. It uses 3 documents:
- A links to B with a weight of 0.6 and to C with a weight of 0.4.
- B links to C with a weight of 0.9.
- C links to A with a weight of 0.5.
The patent notes that "a typical value for a is 0.1" but uses a = 0.5 for the example. The result is r(A) ≈ 0.237, r(B) ≈ 0.202 and r(C) ≈ 0.281. A passes more through its link to B (0.6) than through its link to C (0.4), although both links sit on the same page. C ranks first because it receives 2 links, including B's link weighted 0.9.
Which link features does Google analyze?
The patent lists features of the link itself, of the source document and of the target document, and states that each list "is not exhaustive."
Link features:
- Font size and style: the font size of the anchor text, the font color and attributes such as italics, gray or "same color as background."
- Position on the page: in an HTML list (and the position within the list), in running text, above or below the first screenful viewed on an 800 × 600 browser display, the side of the document (top, bottom, left, right), in a footer or in a sidebar.
- Anchor text: the actual words it contains and its number of words.
- Context: a few words before and after the link, and the topical cluster of the anchor text.
- Link type: for example an image link, and the aspect ratio of the image.
- Commerciality of the anchor text.
- URL features: whether the link leads to the same host or domain, whether the link URL is shorter than the referring URL when both share a domain, and whether the link URL embeds another URL (for example for a server-side redirect).
Source document features:
- URL and site of the source page.
- Number of links on the page.
- Words in the body and in the headings.
- Topical cluster of the source page, and how far it matches the topical cluster of the anchor text.
Target document features:
- URL and site of the target page, and whether it shares a host or domain with the source.
- Words in the URL and the length of the address.
How does the model learn which links get clicked?
The model learns by comparing the links users selected with the links they ignored. The system collects user behavior data: "navigational actions (e.g., what links the users selected, addresses entered by the users, forms completed by the users, etc.), the language of the users, interests of the users, query terms entered by the users." The patent says this data can come from a web browser or a browser assistant, such as a plug-in, that records the documents a user visits and the links selected.
Each selected link counts as a positive instance, and each link left unselected on the same page as a negative instance. When no link on a page is selected, all its links count as negative. The patent's example: document W links to X, Y and Z, and users made the selections (W, X), (W, X) and (W, Z). That gives 3 positive instances and 6 negative ones.
The system then builds a feature vector for each link and trains the model with "a naive bayes, a decision tree, logistic regression, or a hand-tailored approach." Training produces 2 kinds of rules:
- General rules that apply across documents: anchor text above a certain font size raises the selection probability, a link closer to the top of the page is more likely to be selected, and a link between topically related documents is more likely to be followed.
- Document-specific rules: "a link positioned under the 'More Top Stories' heading on the cnn.com web site has a high probability of being selected." Low probabilities go to a link whose target URL contains the word "domainpark," a link on a source document that contains a popup, a link to a target domain that ends in "tv," and a link to a target URL with multiple hyphens.
The patent presents these rules "merely as examples." The model then weighs new links it has never seen clicked: in claim 12, the probability is determined "using only the feature data associated with the second link as input to the model."
Are the link weights fixed?
No, the link weights are not fixed: the patent calls the model "dynamic" because "it is built from data that changes over time." Since links appear and disappear and user behavior keeps changing, the system can periodically update the weights and therefore the ranks.
The weights can also depend on who clicks. The behavior data may cover all users, a class of users or a single user, and the weights "may be tailored to the user class" or to the individual user. The resulting rank is not the whole ranking either: "a document's rank may be one of several factors used to determine an overall rank for the document."
What does the reasonable surfer change for your SEO?
The reasonable surfer means that a link passes authority in proportion to its chance of being clicked. 6 consequences follow:
- Place important links where readers click. The patent's example rule favors links "positioned closer to the top of a document," and placement in running text, a footer or a sidebar is a feature the model weighs.
- Write descriptive anchors that match the topic. The patent weighs the words of the anchor and the match between the topical cluster of the source page and that of the anchor text.
- Link between related pages. Links between topically related documents have a higher selection probability.
- Do not count on links nobody clicks. Terms of Service links, banner ads and links unrelated to the document are the patent's own examples of links rarely followed, and footer or sidebar links can be discounted if users ignore them. Another Google patent goes further and scores quality from links that bring clicks and the traffic they send.
- Keep pages clean of pop-ups and link clutter. The patent's example rule gives a low probability to links on a page with a popup, the number of links on the page is a feature, and in the formula each link's share is divided by that number.
- Never hide links. Gray text and text the "same color as background" are among the link attributes the model reads.
In on-page SEO, we move the links that matter into the running text near the top of the page and remove from templates the links readers never click.
The same logic applies to backlinks. A link buried in the footer of a partner site is likely to pass less than an editorial link that readers click in the middle of an article. Editorial links from unrelated sites also matter to Google Panda, which counts only independent links.
The patent describes what Google's system can do. It does not confirm which features Google weighs today, or how much.