Google predicts the quality of a site it has never scored by comparing the phrases on its pages with the phrases of sites it has already rated. Patent US 9,767,157 B2, "Predicting site quality," describes this phrase model. Google filed it on March 15, 2013, and the patent was granted on September 19, 2017.
What does patent US 9,767,157 describe?
Patent US 9,767,157 describes a method that turns the vocabulary of a website into a predicted quality score. Its 2 inventors are Navneet Panda, the engineer behind the Panda update, and Yun Zhou. The predicted score feeds the ranking engine, which uses it to "rank resources, or to rank search results that identify resources, that are found on one site relative to resources found on another site."
The patent works at site level. It defines a site as "a web site or other collection of data resources," and every computation counts pages across the whole site. The goal is "to predict a score that is comparable to the baseline site quality scores of the previously-scored sites." The patent adds that the same model "can be used to predict the quality of other, smaller collections of text than an entire site, for example, a single page, paragraph, sentence, or even query."
In the claims, the ranking engine uses the baseline score for resources on a previously scored site and the predicted score for resources on a site that has one.
Why does Google need to predict site quality?
Google needs to predict site quality because its regular quality scores are not available for every site. The patent calls them baseline site quality scores and describes them as signals, "among other signals, to rank search results." It explains that they come from "a backend process that may be expensive in terms of time or computing resources, or by a process that may not be applicable to all sites."
The phrase model fills that gap: it can predict a score for a new site, that is a site that is not one of the previously scored sites, "in the absence of other information."
How does Google build the phrase model?
Google builds the phrase model from sites that already have a baseline quality score. Figure 2 of the patent describes the process:
- List the n-grams found on each site of a collection, which can be all the sites indexed by the search system or only the previously scored sites. An n-gram is a sequence of tokens, punctuation included. The patent cites 2-grams and 3-grams, as well as longer 4-grams or 5-grams, of one length or of mixed lengths. The tokenization can be left unnormalized, so it "preserves any errors existing in the original."
- Optionally drop rare n-grams that occur on very few sites, "e.g., fewer than 10, 50, 75 or 100 sites," a threshold "determined based on experience."
- Count the pages of each site that include each n-gram. A page can also be considered to include text from other sources, "for example, from anchor text of links pointing to the pages."
- Compute a relative frequency for each n-gram on each site: the number of pages containing the n-gram divided by the number of pages on the site (or a function of that quotient).
- Map each n-gram to a list of sites, each with its relative frequency for that n-gram.
- Sort the sites into buckets by relative frequency, with "20 to 100" buckets per phrase. The buckets can cover equal frequency intervals, or intervals chosen so that roughly the same number of sites falls into each one.
- Collect baseline scores for the previously scored sites.
- Average the baseline scores of the scored sites in each bucket: an arithmetic or geometric mean, a mode, a median or another measure of central tendency.
The result is a vector for every phrase: one average quality score per frequency bucket. The model can drop neutral phrases whose averages sit "very close" to the global average, the "neutral" site quality score, for example within 0.2% to 5% of it. The patent's example is the 2-gram "on the," which "empirically is almost equally likely to appear in a high quality site as in a low quality site."
What is the relative frequency of a phrase?
The relative frequency of a phrase is the share of a site's pages that contain it. A 3-gram found on 30 pages of a 100-page site has a relative frequency of 0.3. A phrase in a footer repeated on every page reaches 1.0.
The measure counts pages, not occurrences. A phrase repeated 50 times on one page weighs the same as a phrase used once on that page.
How does Google score a new site?
Google scores a new site by looking up each of its phrases in the model and averaging the results. Figure 3 of the patent describes the process:
- Measure the relative frequency of each phrase on the new site.
- Look up the phrase in the model and read the average score of the matching frequency bucket.
- Aggregate the scores of all phrases into one site score: an arithmetic or geometric mean, a mode, a median or another measure of central tendency.
- Smooth the result if needed (an optional step), so that a few phrases cannot dominate the score.
- Predict the site quality score from the aggregate, either directly or through a function of the aggregate.
- Send the predicted score to the ranking engine as the site quality score input.
The patent sets 2 limits. "If the new site has no phrases that are in the phrase model, a prediction cannot be made," and some implementations require a minimum number of phrases found in the model, for example 3, 5, 10, 15, 20 or 50.
How are the phrase scores weighted and smoothed?
The aggregate can be a weighted average. The patent names 3 possible weights:
- the frequency of the phrase on the new site, or a function of it;
- the distance from the neutral score, so phrases far from the global average count more;
- a reduced weight for phrases that occur in only one domain, "representing a judgment that such scores are not fully trustworthy."
Smoothing is triggered by the distribution of the weights. The system divides the sum of the weights of the top N phrases (N being a small number such as 10, 20, 30, 40 or 50) by the sum of the weights of all phrases. If that share is too high, above a threshold such as 0.1, 0.15, 0.20, 0.30 or 0.40, the aggregate is interpolated with the neutral score:
smoothed aggregate score = aggregate score × alpha + neutral score × (1 − alpha)
The coefficient alpha can be 0.90, 0.80, 0.75 or 0.60, or a function of the share, such as a sigmoid: the more the top phrases weigh, the smaller alpha and the stronger the smoothing. The smoothing "can be repeated until an acceptable quotient value is achieved."
Which phrases signal low quality?
The patent publishes no list of good or bad phrases. The model learns them from the data: a phrase becomes a negative signal when the sites that use it at a given frequency have low baseline scores on average.
The same phrase can therefore carry 2 meanings. Used on 5% of a site's pages it can match the profile of good sites, and used on 100% of the pages it can match the profile of weak ones. Frequency, not the phrase alone, carries the signal. Word sequences inside a single page fall under a different method, where Google uses language models to detect spun and stuffed pages.
Is the site quality patent part of Panda?
No, not officially: Google has not stated that this patent is part of Panda. The patent shares its first inventor with Panda and its object, site-level quality, with Panda's stated goal. It was filed in March 2013, 2 years after the first Panda launch and about 6 months after the other site-level patent signed by Navneet Panda.
The 2 patents measure different evidence. The Panda patent on links and brand searches compares independent links with the searches people make for a site. The site quality patent reads the language of the site itself.
What does the site quality patent change for your SEO?
The site quality patent means that the language of your whole site can serve as evidence of its quality, in the absence of other information. 4 consequences follow:
- Write in the vocabulary of the best sources in your field. The model compares your phrase profile with sites already scored, so content that reads like recognized references is more likely to match the profiles of high-quality sites.
- Limit site-wide boilerplate. A sentence in every footer or sidebar reaches a relative frequency of 1.0 and weighs on the phrase profile of the whole domain.
- Avoid templated pages at scale. Hundreds of pages generated from one template repeat the same n-grams at high frequency. If the sites with that profile have low baseline scores, the model carries that average over to yours.
- Launch with complete pages. The prediction is designed for sites without a baseline score: the phrases of your first pages are what the model reads.
An SEO audit from our team measures site-wide boilerplate and templated pages, the 2 sources of repeated n-grams that weigh on the phrase profile of a domain.
The patent describes what Google's system can do. It does not confirm that Google uses this model today, or with these exact parameters.