Subscribe to Newsletter
SpamPatentsContent quality

Gibberish content: how Google uses language models to detect spun and stuffed SEO pages

By 8 min read

Gibberish content is text that is unlikely to be natural language, written to rank for valuable keywords rather than to inform. Patent US 8,554,769 B1, "Identifying gibberish content in resources," describes how Google detects it with a language model and a query stuffing score. Google filed it on June 17, 2009, from a provisional application of June 17, 2008, and the patent was granted on October 8, 2013.

What does patent US 8,554,769 describe?

Patent US 8,554,769 describes a system that computes a gibberish score for each page and uses it to remove or demote the page in search results. Its 4 inventors are Shashidhar A. Thakur, Sushrut Karanjkar, Pavel Levin and Thorsten Brants, a Google researcher on large language models for machine translation whose 2007 paper the patent cites. The patent holds 21 claims and 4 drawing sheets.

The abstract sets out the method: generate "a language model score" by applying a language model to the text, generate "a query stuffing score" as "a function of term frequency in the resource content and a query index," combine them into "a gibberish score," and use it "to determine whether to modify a ranking score of the resource."

What is gibberish content?

Gibberish content is "resource content that is likely to represent spam content." It includes "text sequences that are unlikely, based on specified criteria, to represent natural language text strings (e.g., conversational syntax)," or to represent the non-conversational strings that typically occur in web documents. The patent's example: a spammer can generate a web page that "includes a number of high value keywords such that the search engine will identify the web page as highly relevant."

The patent names 3 ways spammers produce it:

  1. Low-cost labor: text written by "low-cost untrained labor."
  2. Scraping and splicing: "scraping content and modifying and splicing it randomly."
  3. Translation: "translating from a different language."

The spammer earns revenue from the traffic "by including, for example, advertisements, pay-per-click links, and affiliate programs," while the page "typically does not provide any useful information to a user."

How does the language model score work?

The language model score measures how likely each block of text is to be real natural language. The system works in 4 steps:

  1. Parse the page into segments: HTML tags such as headings and paragraphs split the text. Small paragraphs that indicate fragments such as "menu items" can be "removed or otherwise ignored," and sequences of proper nouns (names, geographic locations), which "often represent lists rather than natural language," can be filtered out.
  2. Score each segment with a language model: the system first picks a language model "corresponding to a language of the text segments," and can use several models for segments in different languages. "In some implementations, a 5-gram language model is used," which computes the probability of each word given the words before it. Its example: the likelihood that "sheep" follows "the black." A paragraph can be scored as a whole, or sentence by sentence, with the sentence probabilities summed and divided by the number of sentences.
  3. Normalize by length: for example, each segment's initial score is divided by the paragraph length. "If the resulting text segment score is greater than a threshold value, the text segment is identified as containing gibberish content." Normalization can be skipped when all segments fall within a specified size range.
  4. Compute the page's score: as a function of "the fraction of gibberish terms to the number of terms in the resource" and the sum of the scores of the gibberish segments. Every term of a segment classified as gibberish counts toward the volume of gibberish in the page.

The patent's example of n-gram probabilities is the string "NASA officials say they hope," with n-grams limited to 3 words, scored as a product of conditional probabilities determined "according to relative frequencies in a collection of training data." A smoothing technique (for example back off weights) estimates probabilities for n-grams with sparse training data.

How does the query stuffing score work?

The query stuffing score measures whether a page is packed with real search queries strung together. It relies on a query index built from "a query log identifying queries from a group of users over a specified period of time (e.g., a month)." Particular common queries, punctuation, URLs and other strings that are not natural language can be filtered out. Each term is a key, linked to every query that contains it, and a query can sit under several keys: "cat food" belongs to both "cat" and "food." The patent's example: the key "cat" lists "cat," "cat food," "tabby cat," "cat show breeds," "how to wash a cat" and "funny cat photos."

The system then works in 4 steps:

  1. Find the most frequent terms of the page, for example "the two most frequent, non-filtered, terms," leaving out stop words, very common terms and very rare terms.
  2. Check the phrases around each occurrence against the queries of that term's key. Each matching query is a "hit."
  3. Compute the share of queries hit for each key. "If a particular key is associated with twenty queries, five of which matched phrases in the text content, then ¼ of the queries were hit for that key."
  4. Combine the average and the maximum of those shares into the query stuffing score. With 2 terms, the average is (S1 + S2) / 2 and the maximum is the larger of S1 and S2. With a single term, the score is simply S1. The function can weight one measure over the other, for example as a product of 2 monotonic functions.

The logic: "the more queries identified of all possible queries for the key, the more likely that the text content has been stuffed with queries," because many of them "would be unlikely to occur together in the text content of a single resource."

How does the gibberish score change rankings?

The gibberish score changes rankings through 2 thresholds. In one implementation, the system selects the minimum of the language model score and the query stuffing score as the gibberish score. In other implementations, the language model score boosts the query stuffing score, or either score is used alone. Then:

  1. At or below a first threshold: the page is removed "as a candidate for search results altogether." The system can remove it from the index, move it "to a lower tier index," tag it as unavailable, or apply a very large weight.
  2. Between the 2 thresholds: the page's ranking score is weighted down to demote it. The weight can vary with the gibberish score, for example a factor "in inverse proportion to the gibberish score," so some pages are demoted more than others.
  3. At or above the second threshold: the ranking score stays unchanged.

Can a page flagged as gibberish still appear in results?

Yes, in one case: URL and site queries. The patent describes a filter that "disables weighting of gibberish documents for URL and site queries." If the query is the URL of the page, "the search result for the resource is always returned even if the resource has a gibberish score that is less than the first threshold." A gibberish page can therefore disappear from other queries while staying reachable for anyone who searches its exact address.

The patent states its goal: removing or demoting gibberish "reduces the ability of spammers to receive revenue from generated gibberish content."

Does this patent apply to AI-generated content?

Yes, its logic applies to any text produced at scale without value, even though the patent predates large language models in content production. A language model detects text whose word sequences are unlikely, and a query index detects pages that string together searched phrases. Google's spam policies, since March 2024, name "scaled content abuse": producing many pages, by automation or by people, mainly to manipulate rankings.

Fluent AI text is likely to score well on a language model test, which shifts the burden to other signals, such as information gain and the n-gram profile used to predict site quality.

What does the gibberish patent change for your SEO?

The gibberish patent means that text written for keywords rather than for readers can be measured and filtered. 5 consequences follow:

  1. Write natural sentences. A language model scores the likelihood of word sequences: stuffed, spun or badly translated text scores low.
  2. Do not string queries together. A page that contains a large share of the known queries for its main term looks stuffed, even when each phrase is real.
  3. Edit machine translations. Raw translation is one of the patent's named sources of gibberish.
  4. Keep navigation and lists out of the main text. The patent can ignore short fragments such as menus and filter out lists of proper nouns: the content that counts is your paragraphs.
  5. Publish fewer, better pages. Mass-produced pages built for keywords are the exact target of the patent and of Google's scaled content abuse policy.

An SEO content audit from our team flags the 3 patterns this patent targets: pages that string together query variants, raw machine translations and mass-produced templates.

The patent describes what Google's system can do. It does not confirm that Google uses these exact scores or thresholds today.

Quick quiz

Did you get it?

Test what you just read.

Question 1 of 4

Which 2 scores combine into the gibberish score?

Related articles

We read the patents so you don't have to.

Every week: 3 data-backed SEO insights, 1 myth busted, 1 pattern to steal. No fluff. No guru advice.

Free forever. Unsubscribe anytime. We respect your inbox. Privacy.