Subscribe to Newsletter
Site qualityPatentsPanda

Site quality score: how Google turns brand searches into a quality signal for SEO

By 8 min read

A site quality score is a score computed for a whole site from the ratio between the searches that refer to the site by name and the searches that lead users to click one of its pages. Patent US 9,031,929 B1, "Site quality score," describes it. Google filed it on June 27, 2012, from a provisional application of January 2012, and the patent was granted on May 12, 2015.

What does patent US 9,031,929 describe?

Patent US 9,031,929 describes a site-level score built from user behavior in the search engine, which can be used as a term in the scores of the pages in that site. Its 2 inventors are April R. Lehman and Navneet Panda, the engineer who gave his name to the Panda update. The patent holds 48 claims and 3 drawing sheets.

The abstract defines the method in 3 steps: count the unique queries "categorized as referring to a particular site," count the unique queries "associated with the particular site," and determine "a site quality score for the particular site" from the 2 counts. The patent explains the purpose in one line: "The score is determined from quantities indicating user actions of seeking out and preferring particular sites and the resources found in particular sites."

What is a query that refers to a site?

A query that refers to a site is a search in which the user names or targets that site. The patent recognizes 3 kinds:

  1. Queries with a site label: the "site:" operator followed by a domain name, as in "san francisco site:www.example.com."
  2. Queries with a term that refers to the site: when users commonly type "example sf" or "esf" to mean the site "sf.example.com," the queries "example sf news" and "esf restaurant reviews" count as queries that refer to it.
  3. Navigational queries: queries "submitted in order to get to a single, particular web site." For a search engine, this is "a matter of inference": the patent's example is a query for which a result linked to the site has received "at least a threshold percentage of the user selections" received for all the results shown for that query.

The patent calls them "site queries." The number of unique site queries is S.

What is a query associated with a site?

A query associated with a site is any query followed by a user selection of a result that points to a page of that site. The query "example restaurant reviews" is associated with "example.com" when a user selects a result such as "http://example.com/resource." The topic of the query does not matter. The count of these queries is U.

A selection is not always a simple click. The patent defines it as any action that "causes the resource identified by a search result to be presented, at least in part, to the user." Depending on the configuration, the system can count "a mouse rollover, a click, a click of at least a certain minimum duration," meaning the user views the resource for at least that duration, or a click whose duration is measured "relative to a physical or temporal length of the resource."

How does Google compute the site quality score?

Google computes the score as a ratio with site queries in the numerator and associated queries in the denominator. The base formula is S / U. The patent gives 5 variants:

VariantFormulaEffect
Threshold on site queries(S − T) / USubtracts a threshold T from S, for example a number from 1 to 30 (2, 3, 5, 10, 20 or 30)
Threshold with a floormax(L, S − T) / USame subtraction, but the numerator never drops below a lower bound L, such as zero
DampeningS / U^nRaises U to a power n between 0 and 1, such as 0.5, 0.6, 0.7, 0.75, 0.8 or 0.9
Base value and dampeningS / (B + U^n)Adds a base value B greater than zero, such as 1, 2, 5, 10, 15, 20, 50 or 100
All combinedmax(L, S − T) / (B + U^n)Combines the threshold, the floor, the base value and the dampening

The patent says the threshold "may be determined empirically according to the uses that will be made of the score," and that the power n dampens "the effect of the denominator." It does not explain these choices further. One reading: T keeps a handful of site queries from carrying much weight, B limits the effect of very small denominators, and n keeps very large click volumes from weighing in proportion. The independent claims require at least one of these refinements: a predetermined threshold on the numerator, a power between 0 and 1 on the denominator, or both.

The counts can cover a time window up to the current time, such as "the preceding day, two days, week or month," or all the query data available.

How does Google count unique queries?

Google counts unique queries by grouping repeated queries, with several options the patent leaves open:

  1. Word order: in some implementations, "san francisco site:example.com" and "francisco san site:example.com" count as one unique query. In others, order matters.
  2. Site label: the placement of the site label within the query can optionally be ignored.
  3. Users: the same query can count as 2 unique queries when 2 different users submit it. A user can be identified by a logged-in account, an IP address or an Internet cookie, and the patent states that this information "can be anonymized," for example with quasi-unique identifiers instead of the user's actual identity.

What counts as a site?

A site is a collection of resources defined in one of 4 ways:

  1. A server: the resources hosted on a particular server.
  2. A domain: "example.com" and every host under it.
  3. A subdomain: "www.example.com" and its pages.
  4. A subdirectory: "example.com/subdirectory" and its pages.

The patent leaves this choice to the configuration of the system. It matters for multi-section sites: depending on that configuration, a blog in a subdirectory can share the score of the whole domain or receive a score of its own.

How does the score affect rankings?

The score affects rankings as a term in the computation of scores for the pages of the site. The patent's example: the site quality score of "http://www.example.com" can be "used as a term in the computation of a score for a resource 'http://www.example.com/resource.html'." The patent gives 2 other uses:

  1. Ranking one site against another: the score can rank resources, or results that identify them, "that are found in one site relative to resources found in another."
  2. Weighting other signals: "A high site quality score for a particular site can be used to determine how other attributes are used to score resources in that site."

Can the score be computed from clicks instead of queries?

Yes. A second process, claimed on its own, replaces counts of unique queries with counts of user selections:

  1. Numerator: the number of selections on results shown for queries that refer to the site. With "example sf" as a term for "sf.example.com," the selections made after "example sf news" or "example sf restaurant reviews" count.
  2. Denominator: the number of selections on results that identify pages of the site, "regardless of the content of the query."

The same 5 ratios apply, with the same threshold, floor, base value and power. In this version, every click counts, not only unique queries. Selections combined with links are the basis of another patent, which scores quality from links and the traffic they send.

How does this patent relate to Panda?

This patent shares Panda's first premise: quality is judged at the level of the site, from the behavior of users. Navneet Panda co-signed it, and it was filed in 2012, the same year as the Panda patent on links and brand searches.

The 2 patents place brand searches on opposite sides of their ratio. In the Panda patent, searches that refer to the site sit in the denominator: brand demand without independent links lowers the factor. In the site quality score, the same searches sit in the numerator: people who seek the site out by name raise its score. One reading of the 2 patents together: brand demand helps a site most when independent sites also cite it.

The site quality prediction patent, also signed by Navneet Panda, takes over when such behavior data is missing, for new sites.

What does the site quality score change for your SEO?

The site quality score means that people looking for your site by name are evidence of its quality. 5 consequences follow:

  1. Build a name people search for. Navigational and brand queries feed the numerator. Every channel that makes people search for your brand (newsletter, video, social media, events) feeds it.
  2. Make your brand easy to search. A short, distinctive name that users can type, and that Google can associate with your site, turns mentions into site queries.
  3. Win the navigational click. The patent's example infers that a query is navigational when your result receives at least a threshold share of the selections for it. Own the results for your own name, with a clear homepage title and sitelinks.
  4. Do not chase traffic alone. Clicks from generic queries grow the denominator. Traffic without brand demand lowers the ratio, even if the dampening power softens the effect.
  5. Choose your site structure on purpose. The system can be configured to score a site by server, domain, subdomain or subdirectory: depending on that configuration, a weak section in its own subdomain can be judged apart from the main site, and vice versa.

The patent describes what Google's system can do. It does not confirm that Google uses this ratio today, or with these parameters.

Quick quiz

Did you get it?

Test what you just read.

Question 1 of 4

What sits in the numerator of the site quality score?

Primary source for this article:

US 9,031,929 B1: Site quality score

Related articles

We read the patents so you don't have to.

Every week: 3 data-backed SEO insights, 1 myth busted, 1 pattern to steal. No fluff. No guru advice.

Free forever. Unsubscribe anytime. We respect your inbox. Privacy.