Subscribe to Newsletter
User signalsPatentsClicks

Presentation bias in SEO: how Google discounts clicks you only got from position and bold text

By 11 min read

Presentation bias is the share of clicks a search result receives because of how it is displayed, independently of its quality. Patent US 8,938,463 B1, "Modifying search result ranking based on implicit user feedback and a model of presentation bias," describes how Google models that bias and removes it from click signals. Google filed it on March 12, 2007, and the patent was granted on January 20, 2015.

What does patent US 8,938,463 describe?

Patent US 8,938,463 describes a prior model that predicts how often a result should be clicked given how it is presented, used to correct click-based ranking. Its 2 inventors are Hyung-Jin Kim, an inventor of Google's click patent, and Adrian D. Corduneanu. The patent holds 27 claims and 8 drawing sheets, and its term was extended by 836 days.

The abstract defines the method: receive features, including "a first feature indicative of presentation bias," collect "document result selections," generate "a prior model" that represents "a background probability of document result selection given values of the multiple features," and output it "to a ranking engine for ranking of search results to reduce influence of the presentation bias."

Why do clicks need correcting?

Clicks need correcting because users click results for reasons unrelated to quality. The patent states that "users tend to click results with good snippets, or that are higher in the ranking, regardless of the real relevance of the document to the query." It names "an attractive title or snippet" and the position of the result. Its example: "the top result will receive many more clicks than the last result, even if their positions are swapped." Results are also spread over several pages, so "many of the document results for a typical search may never be viewed at all by the user."

The patent reuses the click framework of Navboost: clicks are classified by duration, with example weights of −0.1 for a short click, 0.5 for a medium click, 1.0 for a long click and 0.9 for a last click (only 0.3 when another click preceded the last one). These weighted clicks feed the "traditional click fraction," BASE = #WC(Q,D) / (#WC(Q) + S0): the weighted clicks on a result for a query, divided by the weighted clicks on all results for that query plus a smoothing factor. Language and country versions fall back to the broader fraction when data is scarce.

Which features create presentation bias?

The patent's table lists 19 features that can raise or lower click rates, often independently of quality. The main ones:

FeatureEffect on clicks described in the patent
PositionResults in higher positions can get higher click-through rates
PageResults in higher pages can get higher click-through rates
Bold termsResults with more bolded terms in the title, snippet and URL can get higher click-through rates
Position of bold termsBold terms have a different impact at the beginning or at the end of the sentence
Length of title, snippet and URLResults with longer titles, snippets and URLs can get higher click-through rates
Attractive termsSexually suggestive terms, or more broadly "any term that increases the click-through rate of any result in which it appears, independent of the actual query or url"
Quality of neighboring resultsHigh-quality results in previous or next positions can lower a result's clicks
Previous and next pagesUsers' attraction to previous or next pages can affect a result's clicks
AdsBetter quality ads can lower the click rate of all results
OneboxSpecial results shown before normal results (maps, book search) can lower their clicks, "especially if they are very topical"
Query lengthGood results may be lacking for longer queries
Query type and queryQuery clusters and specific queries have different click-through rates
Format unrelated to rankIndentation (for example a deeper page of the same site) can raise or lower clicks
Format linked to rankA light blue background on some top results can raise or lower clicks
Other UI featuresAnything that makes a result stand out in the list can affect its clicks

The patent adds the features of custom search engines on external sites, such as "the layout of the page, the color of the links, flash ads on the page."

Is position a bias or a quality signal?

Position is both. The patent calls it "really a mixed feature since clicks based on position are influenced by both the presentation bias aspect and the result quality aspect." In some implementations, position is therefore treated as a quality feature, and the bias features are limited to those "unconnected with result quality," such as bolding and the length of the title and snippet. The claims, however, require the rank of the result to be one of the presentation bias features.

In general, the features chosen are "correlated with click rate" and "independent of the query (Q) and the search result (D)." In some implementations the specific query, or its type defined by query clustering, can still be used as a feature.

How does Google build the prior model?

Google builds the prior model from historical click logs, across queries. For each combination of feature values, the system counts clicks of each kind, "separately accumulating short, medium and long clicks, as well as the event of a result being shown, but another result being clicked." The patent's example: "the number of long, short and medium clicks at position 1 on an English language search site," along with the number of times other results were clicked while a result was shown at position 1. The log entry of a clicked result can also store features of the results returned with it, such as their IR score, even if they were never clicked.

The model can be "a standard linear or logistic regression, or other model types," and it "can essentially represent the statistics of historical click data, indicating what percentage of people clicked on a result given presentation of the result and the set of features." When there are too many combinations of feature values, the features are split into 2 disjoint groups: click counts are accumulated for every combination of a reduced set of more important features, and the model is trained on the values of the remaining features. This allows more parameters "but not so many parameters that the model cannot be trained."

How does the two-model variant isolate the bias?

The two-model variant compares a model built on quality features with a model that adds presentation features. In the patent's example, the quality model uses position, IR score, the IR score of the top result, the IR scores of the previous and next results, country, language and the number of words in the query. The second model adds the number of bold terms in the titles and snippets of the current, previous and next results, the lengths of those titles and snippets, and Boolean indicators of pornographic terms, of an ad and of a special result. "Any difference between these two models then is likely to be due only to presentation bias."

For every result, the ranking score is multiplied by the ratio of the click fraction predicted by the quality model over the one predicted by the full model, or by a monotonic function of that ratio. The rule behind it: "if a result is expected to have a higher click rate due to presentation bias, this result's click evidence should be discounted; and if the result is expected to have a lower click rate due to presentation bias, this result's click evidence should be over-counted."

How does the model change rankings?

The model changes rankings by comparing the actual click fraction of a result with the click fraction predicted from its presentation. For a search, the feature values of each result are looked up in the prior model, and the predicted click-through rate is compared with the traditional click fraction. In the position example, the ranking score "can be multiplied by a number smaller than one if the traditional click fraction is smaller than the click fraction predicted by the prior model, or by a number greater than one otherwise."

The patent gives 3 example transforms, where a is the actual click fraction, p the predicted one and the other letters tuned parameters:

  • Boost = C(a/p)
  • Boost = max(min(1 + Z, m0), m1), with Z = (a - p)^k1 if a ≥ p and Z = -1 * abs((a - p)^k2) if a < p
  • the same bounded boost, with Z = C[1/(1 + e^(-k(a - p)))] - 1/2

The transform is chosen by tuning on historical data "combined with human generated relevance ratings." The correction can be applied at run time, after the clicks are aggregated, or at model building time, where each click can be weighted "according to how much they were affected by display bias before the clicks are aggregated." The patent notes that the run time adjustment "can result in some loss."

How are results without click history scored?

Results without click history receive a click fraction predicted by a separate prior model. To build it, the logs are randomly split into 2 halves, and the results that have click evidence in the first half but none in the second are extracted: they represent the less frequently selected documents. A prior model is trained on them with features such as position, country, language and IR score. A new result whose traditional click fraction is undefined, "because the result has not been selected before," is then assigned the click fraction predicted by this model.

The patent also describes combining a relevance signal (such as the click fraction) with a presentation signal from a prior model into a relevance signal "generally independent of presentation," which is then sent to the ranking engine.

What do the claims protect?

The claims protect a model trained to predict a click-through rate from presentation bias features and relevancy features, used by a search engine to determine a quality score for each result and factor out "independent effects of presentation bias." The 3 independent claims (a method, a system and a storage device) require at least one presentation bias feature to be the rank of the result. Dependent claims list as bias features the position of another result, the title, its length, the snippet, bolded text in the snippet and the presence of ads; and as relevancy features the IR score of the result, the IR score of another result, the language of the query and the number of words in the query. Other dependent claims cover the two-model ratio and the comparison with the click fraction, "a ratio" or "a value difference."

How does this patent relate to Navboost?

This patent complements the click system described in Navboost. Navboost's patent compares a result with itself to resist presentation bias; this patent models the bias explicitly and factors it out of the click signal. Both share an inventor, Hyung-Jin Kim, and the same example click weights. The patent adds that other click models can replace the traditional click fraction, such as "a large-scale logistic regression model that uses the actual query and url as features."

What does presentation bias change for your SEO?

Presentation bias means that clicks you win only through position or eye-catching formatting are discounted, while clicks you win beyond expectations count more. 6 consequences follow:

  1. Beat the expected click rate of your presentation. In the patent's examples, a result clicked more than its position and display predict gets a boost, and a result clicked less gets a multiplier below one.
  2. Do not rely on clickbait. Attractive terms that raise clicks regardless of the query are modeled as bias, and the click evidence they generate is discounted.
  3. Earn long clicks, not just clicks. The model counts short, medium and long clicks separately, and a short click carries a negative weight (−0.1) in the patent's example.
  4. Do not count on bold terms and long titles alone. The patent models bolded terms and the length of titles and snippets as presentation features: the clicks they attract are expected, so the page has to deliver more than the snippet promises.
  5. Know your competition on the page. Strong neighboring results, ads and special results lower your expected clicks: the model accounts for them, so you are judged against realistic expectations.
  6. A new page is not judged on an empty record. A result with no click history can receive a click fraction predicted from features such as its position, the country, the language and its IR score.

An SEO audit from our team compares the click rate of each page in Search Console with the rate expected at its position, to find pages clicked less than their position predicts.

The patent describes what Google's system can do. It does not confirm which bias features Google models today.

Quick quiz

Did you get it?

Test what you just read.

Question 1 of 4

What does the prior model represent?

Related articles

We read the patents so you don't have to.

Every week: 3 data-backed SEO insights, 1 myth busted, 1 pattern to steal. No fluff. No guru advice.

Free forever. Unsubscribe anytime. We respect your inbox. Privacy.