Subscribe to Newsletter
User signalsPatentsContent updates

Page rewrites, click history and SEO: how Google weighs past clicks when a page changes

By 8 min read

Version-weighted click statistics are click-based quality signals that Google keeps for each version of a page and weights by how much the content has changed since that version. Patent US 9,002,867 B1, "Modifying ranking data based on document changes," describes them. Google filed it on December 30, 2010, and the patent was granted on April 7, 2015.

What does patent US 9,002,867 describe?

Patent US 9,002,867 describes a way to keep using past click data for a page after its content changes, in proportion to how similar the old and new versions are. Its 2 inventors are Henele I. Adams and Hyung-Jin Kim, an inventor of Google's click patent. The patent holds 39 claims and 6 drawing sheets.

The abstract defines the core: "each quality of result statistic is weighted by a weight determined from at least a difference between content of a reference version of the document and content of the version of the document corresponding to the version specific quality of result statistic." The 3 independent claims (a method, a system and a computer-readable medium) all require one specific way of measuring that difference: comparing the time distributions of shingles of the old and current versions, described below.

Which problem does the patent solve?

The patent solves the mismatch between old clicks and new content. Quality of result statistics come from past user behavior, for example "how frequently a user selected a search result corresponding to the given document." When a page changes, clicks earned by the old content may say nothing about the new one.

The patent's drawing shows a personal home page that first talks about animals, with clicks from queries like "dog breeds," "wombat facts" and "information on cats," then switches to desserts: chocolate éclairs, brownies, cupcake frosting. Clicks from "wombat facts" no longer describe the page. The patent's answer: "rather than ignoring all past quality of result statistics when a document changes, a search system can weight the quality of result statistics, for example, by a score derived from how much the document has changed."

How does Google measure how much a page changed?

The system measures change by comparing each past version with a reference version, in some implementations "the most recent version of the document obtained by the search system," such as "the latest version of the document obtained during a crawl." The versions are stored at "a same address," the same URL. The result is a difference score, and the patent describes 4 ways to compute it:

  1. Shingles: "contiguous subsequences of tokens in a document," extracted from the page or from its snippets. The system compares the 2 sets of shingles to get a similarity score, then takes its inverse as the difference score. The example formula is the number of shared shingles divided by the number of distinct shingles across both versions. A variant divides by the shingles of one version only, which "gives greater importance of the changes to the newer version," and shingles can also be weighted by their inverse document frequency.
  2. The time distribution of shingles: each shingle is tied to the first time it was observed, and the distributions of those times are compared, for example through "the distance between the mean of one distribution and the mean of the other distribution." Similar versions have close means; a version that changed dramatically does not.
  3. Text comparison: on the whole text or on the parts significant for ranking, for example by finding "the longest common subsequence of the two versions" and computing "the percentage of the two versions that overlap."
  4. Fingerprints: a hash of the page's shingles. 2 fingerprints are compared bit by bit (an exclusive or, then a sum of the differing bits), with an extra penalty when several fingerprints only match out of order.

The comparison can use only the non-boilerplate text. Boilerplate is identified by comparing related documents, for example from the same domain, and keeping "text that is common to all, or a majority, of the documents," such as common text on the left side or at the bottom. Stripping it means a change of template need not count as a change of content.

How are past clicks weighted?

Past clicks are weighted by the similarity between their version and the reference version. Each version gets its own quality of result statistic for a query; the system weights each one by a weight derived from the difference score and combines them into a "weighted overall quality of result statistic." The rule is explicit: "difference scores indicating larger differences should result in smaller weights than difference scores indicating smaller differences." A version close to the current page keeps more weight than a version that bears no resemblance.

The difference score is not the only factor. The weight can also depend on:

  1. The number of times the page changed since that version was detected: "the larger the number of times the document changed since the version was detected, the lower the weight." A page edited often therefore discounts its old versions faster.
  2. How long later versions stayed unchanged: the longer a later version stayed stable, the lower the weight of the old one.
  3. How much data was collected for later versions: the more data the newer versions have, the less the old version's data counts.

The function that turns these factors into a weight can be linear, quadratic, a step function or a polynomial, "hand-tuned manually" or derived by machine learning. The statistics can be specific to a geographic location, a language, or both.

The system stores the weighted statistic for each query and document pair, and in some implementations a non-weighted one as well, combined without the difference-based weights. When a crawl detects that the page has changed, the system can recompute the difference scores against the new version and re-weight the statistics.

When does Google use the weighted or the non-weighted statistic?

The system chooses with factors compared to thresholds, or combined into one score compared to a threshold. The patent names 4 factors:

  1. Recent change: if the most recently crawled version differs enough from the version used when the weighted statistic was computed, the system uses the non-weighted statistic, because the weights "do not accurately reflect the most recently crawled version of the document." The threshold can be set empirically.
  2. The page's own turnover: some documents, "for example, the home page of a news website, have frequent turnover in content." Above a threshold, the system uses the non-weighted statistic, because the weighted one "reflects a single moment in the frequently changing history of the document."
  3. Turnover of the competing pages: the system takes the top documents for the query, counts those that change more often than a first threshold, and if that count exceeds a second threshold, uses the non-weighted statistic for the query.
  4. Query category: categories such as sports, celebrities or commercial queries can be tied to one statistic or the other. Queries categorized as "seeking recent information" use the weighted statistic, and the patent's example is celebrities, "because these queries are more likely to be seeking the latest information."

The patent adds a date cutoff. If information relevant to a query "would not have existed prior to a particular date," the system can recompute the weighted statistic to minimize data collected before that date, for example by giving a zero weight to older versions. Its example is the query "Results of Summer Olympics 2008," for which data from before 2008 is not relevant.

Can a page be penalized for not changing?

Yes: in some implementations, the patent describes a penalty for pages that stay static while competing pages change. The system "penalizes the weighted or non-weighted overall quality of result statistic for a given document and query when the given document does not change very much over time and other documents responsive to the given query," such as those with the highest statistics, "do change over time." Change is measured by its frequency or by the amount of content that changes. The page's change counts as low when it falls below a threshold computed from the change of the other documents, and the penalty consists, for example, in "reducing the value of the statistic."

The same idea of freshness relative to the query appears in the link velocity patent, where a falling rate of new links marks a page as stale.

What does this patent change for your SEO?

The patent means that a page carries its click history through small updates, but loses it in proportion to how much a rewrite changes it. 6 consequences follow:

  1. Update a page that performs, do not replace it. Small, focused updates keep the page close to the versions that earned its clicks, so their weight is preserved.
  2. Plan big rewrites carefully. A page rewritten from scratch keeps little of its click-based history: expect to rebuild its signals.
  3. Do not repurpose a URL for an unrelated topic. Like the home page that moved from wombats to brownies, old clicks stop describing the page, and the new topic has to earn its own.
  4. Template changes are not content changes. Boilerplate shared across the site can be excluded from the comparison: a redesign alone need not reset history.
  5. Do not edit for the sake of editing. Each change recorded since a version was detected can lower that version's weight, so a stream of minor edits can erode old click history faster than one meaningful update.
  6. Keep pages fresh where competitors are fresh. On queries where the top pages change often, a page that never changes can be penalized. Google can tell which queries call for recent pages from the queries themselves, as the patent on query freshness describes.

Our content marketing work follows rule 1: pages that earn clicks get focused updates, and full rewrites are kept for pages with little click history.

The patent describes what Google's system can do. It does not confirm how Google weighs past clicks after a page change today.

Quick quiz

Did you get it?

Test what you just read.

Question 1 of 4

How is each version's click statistic weighted?

Related articles

We read the patents so you don't have to.

Every week: 3 data-backed SEO insights, 1 myth busted, 1 pattern to steal. No fluff. No guru advice.

Free forever. Unsubscribe anytime. We respect your inbox. Privacy.