Subscribe to Newsletter
PandaPatentsSite quality

The Google Panda patent: why your SEO links must keep pace with your brand searches

By 9 min read

The Panda patent scores a whole site, not a single page, by comparing the independent links it earns with the searches people make for it. Google was granted patent US 8,682,892 B1, "Ranking search results," on March 25, 2014. Its first inventor, Navneet Panda, is the engineer who gave his name to Google's Panda update.

What is the Panda update?

The Panda update is a Google ranking change that demotes low-quality sites as a whole. Google launched it in February 2011, and it changed results for 11.8% of queries in the United States. In January 2016, Google confirmed that Panda had become part of its core ranking algorithm.

Amit Singhal and Matt Cutts explained the name in a March 2011 interview with Wired: Navneet Panda is the Google engineer whose work made the update possible.

The patent was filed on September 28, 2012, 19 months after the first Panda launch. Google never stated that this patent is the Panda algorithm. It describes a site-level quality method signed by the same engineer, which is why the SEO industry calls it "the Panda patent."

How did Panda evolve between 2011 and 2016?

Panda evolved from a periodic filter into a part of the core algorithm in 5 years. The 4 milestones are:

  1. February 2011: the first launch, in English in the United States, targets content farms and shallow pages.
  2. May 2014: Panda 4.0, a major refresh announced by Matt Cutts.
  3. July 2015: Panda 4.2, the last confirmed refresh, with a rollout spread over several months.
  4. January 2016: Google confirms that Panda runs inside the core ranking algorithm and no longer ships as separate updates.

Between refreshes, a demoted site stayed demoted until the next run, even after a cleanup. Since 2016, Panda is part of the core algorithm, but Google's Gary Illyes specified that it does not run in real time.

What does Panda judge on a site?

Panda judges the overall quality of a site's content, from the point of view of a reader. In May 2011, Amit Singhal published 23 questions on the Google Webmaster Central blog to describe what Google means by a high-quality site. 7 of them sum up the logic:

  1. Trust: "Would you trust the information presented in this article?"
  2. Expertise: is the article written by an expert or enthusiast who knows the topic well, or is it shallow?
  3. Redundancy: "Does the site have duplicate, overlapping, or redundant articles on the same or similar topics with slightly different keyword variations?"
  4. Originality: "Does the article provide original content or information, original reporting, original research, or original analysis?"
  5. Intent: are the topics driven by the genuine interests of readers, or by guesses about what might rank in search engines?
  6. Editorial care: "Does this article have spelling, stylistic, or factual errors?"
  7. Ads: "Does the article have an excessive amount of ads that distract from or interfere with the main content?"

The same post states the site-wide effect plainly: "low-quality content on some parts of a website can impact the whole site's rankings." The patent's group-level factor follows the same logic: one score for the domain, applied to every page.

What does patent US 8,682,892 describe?

Patent US 8,682,892 describes a modification factor computed for each group of resources and applied to the score of every page in the group. In the patent's examples, a group is every page under one domain name (example.com/*) or one host name (blog.example.com/*), and pages are grouped so that a resource "cannot be included in more than one group of resources." The 2 inventors are Navneet Panda and Vladimir Ofitserov. The patent was filed on September 28, 2012 and contains 46 claims and 5 drawing sheets.

The patent states its goal plainly: "Search results identifying low-quality resources can be demoted in a presentation order of search results."

The method works in 4 steps:

  1. Count the independent links pointing to the group.
  2. Count the reference queries that target the group.
  3. Combine the 2 counts into the group's modification factor. In the patent's example, the factor is links divided by queries.
  4. Apply the factor to the initial score of every page in the group for future searches. The initial score can be "a measure of the relevance of the resource to the search query, a measure of the quality of the resource, or both," and the factor is typically multiplied with it.

An independent link is a link from a resource that has no relation to the target site. The patent excludes 3 kinds of links from the count:

  • Internal links: the source and the target sit in the same group.
  • Links between related sites: groups "owned by the same entity, hosted by the same entity, or that were created by the same entity."
  • Links between look-alike pages: identical or similar content, images or formatting, down to "identical or similar Cascading Style Sheets (CSS)."

The patent describes 2 ways to decide. In some implementations, every factor must indicate that the 2 pages are independent. In others, the system computes an independence score and checks it against a criterion.

In some implementations, the system counts at most one link per source group, so 500 links from one partner site weigh as much as 1. Alternatively, the count can be the total number of links from that source group, its logarithm "or other non-decreasing function of the actual number." The patent's definition of a link covers implied links too: "a reference to a target resource, e.g., a citation to the target resource" that is not a hyperlink a user can follow. The same one-per-host rule limits LocalRank links, which come from pages that already rank for the query.

What is a reference query?

A reference query is a search that refers to the site itself. The patent describes 2 types: queries that contain a term pointing to the site, such as its domain name, and navigational queries, submitted "in order to get to a single, particular web site or web page of a particular entity." The term does not have to be the URL: if users commonly call sf.example.com "example sf" or "esf," the queries "example sf news" and "esf restaurant reviews" count for that group.

In some implementations, the system counts only queries from unique users, identified by a cookie or a login identifier. The count can cover a specified time period or every reference query on record.

Reference queries measure demand for the site. These navigational searches are the same query type that Google's click patent handles with its own click thresholds.

How does the ratio change rankings?

In the patent's example, the modification factor is M = IL / RQ: independent links divided by reference queries. It follows from the ratio that a site many people search for but few independent sites cite gets a low factor.

The patent does not apply M directly. For each page in the results, it first checks the query, then the page's initial score (IS):

  1. The query is navigational to the page: the factor is set so that it does not change the score (a multiplicative factor of 1).
  2. Below a first threshold T1: the score stays unchanged.
  3. Between T1 and a second, higher threshold T2: the factor is printed as f1 = T1 + (IS - T1) · M / IS, and the patent says the system "applies a modification factor that decreases as the initial score increases."
  4. Above T2: a second factor f2 is computed from f1, with a logarithm of the initial score in base T2 and a smoothing function g(f1). When f1 exceeds a predetermined threshold Q, the factor has "a muted effect or no effect" on the score.

How does the patent normalize the factor?

The patent normalizes the factor within groups that receive similar numbers of reference queries, "because the modification factors may not scale well as the counts of reference queries increase." The system splits the groups into partitions by range of reference queries, computes a statistical measure m of the factors in each partition (a mean, the median, the mode, or a maximum or minimum), then stores a normalized factor NM = (M - m) / m. A small site is then compared with sites of similar demand, not with the most searched sites on the web.

What does the Panda patent change for your SEO?

The Panda patent rewards sites whose reputation outside their own walls matches their popularity. 5 consequences follow:

  1. Earn links from unrelated sites. Links from your other domains, your host's network or templated sister sites do not count as independent.
  2. Grow citations with brand demand. A brand that grows in searches without growing in citations lowers its own ratio. Pair every visibility campaign with coverage from independent sources.
  3. Diversify sources over volume. In one implementation, only one link per source group counts, and in the others extra links from the same source add less and less, so 50 sites citing you once beat 1 site citing you 500 times.
  4. Think at domain level. The factor applies to every page in the group: one domain carries one ratio, and a weak ratio limits even your best page.
  5. Expect the effect on generic queries, not on your name. When a query is navigational to your page, the patent leaves the score unchanged, and pages under the first threshold are not touched either. The factor weighs on the queries where you compete with other sites.

How do you recover from a Panda demotion?

Recovery from a Panda demotion means raising the average quality of the whole domain, not fixing one page. Google's May 2011 guidance names 3 options for weak pages: "removing low quality pages, merging or improving the content of individual shallow pages into more useful pages, or moving low quality pages to a different domain." A recovery plan runs in 4 steps:

  1. Inventory every indexed URL with its organic clicks, links and word count over the last 12 months.
  2. Remove or noindex the pages that bring no traffic, no links and no value to a reader.
  3. Merge the overlapping pages that target the same topic with keyword variations, and redirect them to one complete page.
  4. Rewrite the pages you keep against the 23 questions: sourced facts, original analysis, no spelling errors, ads that never cover the main content.

A technical audit from our SEO agency runs step 1 on the full index: every indexed URL with its clicks, links and word count, so the removal and merge lists come from data.

A recovery can take months, because Google needs time to recrawl and reassess the whole domain.

The patent describes one site-level quality method. Google confirms that Panda exists and runs inside its core algorithm, not that it uses this exact ratio.

Quick quiz

Did you get it?

Test what you just read.

Question 1 of 4

What is the modification factor of the Panda patent?

Primary source for this article:

US 8,682,892 B1: Ranking search results

Related articles

We read the patents so you don't have to.

Every week: 3 data-backed SEO insights, 1 myth busted, 1 pattern to steal. No fluff. No guru advice.

Free forever. Unsubscribe anytime. We respect your inbox. Privacy.