Temporal score adjustments are changes to a page's ranking score based on whether the queries that lead to it ask for new or old information, and on when those queries become popular. Patent US 8,924,379 B1, "Temporal-based score adjustments," describes them. Google filed it on March 5, 2010, and the patent was granted on December 30, 2014.
What does patent US 8,924,379 describe?
Patent US 8,924,379 describes 3 methods that use the timing and wording of queries to adjust rankings. Its 2 inventors are Hyung-Jin Kim, an inventor of the click re-ranking patent, and Andrei Lopatenko. The patent holds 9 claims and 13 drawing sheets.
The abstract lists the 3 methods:
- Age classification: "adjusting the scores for the documents according to an age classification for one or more related queries for the documents."
- Time trends of extended queries: "storing time trend data associating the query and one or more periods of time with a respective extended query."
- Popularity change windows: adjusting scores "according to whether the first time is within a popularity change time window for one or more related queries."
How does Google classify a page as new or old?
Google classifies a page from the temporal terms of the queries that lead to it. Each page is associated with related queries: past queries for which it was a search result. In the description, a query is associated with a page when a user clicked the page's result for that query; in some implementations the user must also stay on the page beyond a dwell time threshold, or the page must have been selected a threshold number of times or by a threshold number of users. In the claims, the related queries must also match the query the user submits. The patent's example: for the query "Super Race Results," the age of a page about the 1996 race is judged only from "Super Race Results 1996," "Old Super Race Results" and "Super Race Results," not from its other related queries.
A temporal term "conveys information about the time reference of the query," whether it asks "about new information or old information." The system compares terms with a list of new temporal terms and a list of old ones, which "can include dates and terms that connote the appropriate temporal meaning." Old temporal terms include old dates ("1996" when the year is 2010) and words such as "old," "previous" or "last year's." New temporal terms include current dates ("2010" when the year is 2010) and words such as "new," "current," "future" or "today." In some implementations, a temporal term already present in the user's query is ignored: for the query "new marathon results," the word "new" in a related query does not count.
Each related query is classified as new, old or non-temporal. The patent's example is Document A with 4 related queries:
| Related query | Classification |
|---|---|
| Marathon 2010 | New |
| Current Marathon | New |
| Marathon 1998 | Old |
| Marathon Results | Non-temporal |
With an unweighted count, the new count is 2 and the old count is 1. The page is classified as new if the new count satisfies a threshold, and otherwise as old if the old count does; a page that is neither is non-temporal. Some implementations add further conditions, such as the new count divided by the total number of related queries exceeding another threshold, or the old count staying below a threshold. A query that contains both new and old temporal terms can be classified as old, new or non-temporal depending on the implementation.
The counts can be weighted by a quality of result statistic for the page and the query: "each query can be weighted by the number of long clicks for the query and the document divided by the total number of long clicks for the query and the document." In the weighted version of the example, the 2 new queries have statistics of 0.8 and 0.7, so the new count is 1.5; the old query has 0.01, so the old count is 0.01. With a new count threshold of 1.4, Document A is classified as new.
How does the age classification change rankings?
The age classification changes rankings with a positive or negative adjustment. The adjustment "increases the score by a first predetermined factor when the document is a new document and decreases the score by a second predetermined factor when the document is an old document." A non-temporal page receives no adjustment. Each factor can be a fixed amount added to or subtracted from the score, or a fixed amount the score is multiplied by. In the claims, the boost factor is determined from the new count and the penalty factor from the old count: the more new queries a page has, the larger its boost can be. In some implementations, the factors are chosen so that every new page scores higher than every old page, while a new page does not overtake another new page that initially had a higher score.
In some implementations, the system first checks that the user's own query is not an "old query": "If the query is an old query, then the user may not be particularly interested in more recent documents," and no adjustment is made. The patent's starting assumption is that "unless queries explicitly suggest otherwise, a user is looking for the most recent information for the subject of their query."
What are extended queries and time trends?
Extended queries are queries that contain every term of a base query plus one or more additional terms. The method starts from recurring queries, queries with multiple spikes in popularity over a period such as a year or a month. The patent's example is "playoff schedule," which spikes in January, April and October, the playoff seasons of the National Football League, the National Basketball Association and Major League Baseball. Its extended queries can be "playoff schedule nfl," "playoff schedule nba" and "playoff schedule mlb."
In some implementations, the extended queries are query refinements: queries a user submits after the base query. A refinement may have to include one or more terms of the base query: "Olympics 2010" is a refinement of "Olympics," but "Winter competition 2010" is not. It may also have to follow the base query within a threshold number of queries or amount of time, for example "five seconds later," and in some implementations the user must not have selected any result for the base query in between. Reformulations within a session feed another Google method, in which contextual synonyms are learned from users who change a single phrase of their query.
For each period of time (the patent's example is two weeks), the system computes a popularity score for each extended query, for example the number of times it "was submitted as a query refinement during the first period," divided by the number of times the base query "was submitted during the first period." It stores, for each period, the extended query with the highest score; some implementations store several extended queries whose score exceeds a threshold. When a user searches the base query during that period, documents can be scored "based, at least in part, on the first extended query." In practice, the ranking engine can give higher scores to documents that contain terms of the extended query absent from the original query, or to documents associated with queries that contain those terms, for example by a factor derived from the extended query's popularity score for the period.
What is a popularity change time window?
A popularity change time window is "a reoccurring period of time during which a popularity of the query temporarily changes beyond a threshold amount." The patent illustrates it with the query "how to cook a turkey": its popularity stays relatively constant until just before November, rises sharply, and returns to its previous level just after the beginning of December, a spike the patent attributes to users in the United States preparing Thanksgiving. Another example is "Easter Bunny," which may become more popular around Easter each year.
The threshold can be set empirically "so that small background fluctuations in popularity are not identified" as windows, and spikes or dips can be detected with conventional time series analysis. The system can aggregate several past years, for example day by day across the calendar year. A query falls within a window when its submission time corresponds to it: if the window is January, based on data from January 2000 to January 2009, a query submitted on January 2, 2010 falls within it.
When a query is submitted during a spike for a page's related queries, the page's score can increase "by a first predetermined factor"; during a dip, it can decrease "by a second predetermined factor." If no related query of the page is in a window, the page receives no adjustment. The factor can also vary: it can be larger when the window is shorter, when the spike or dip is larger, or when more of the page's related queries are in a window. One variant adjusts every page, weighting each related query by its popularity in the current period divided by its average popularity. Seasonal pages can therefore rise when their season comes back. The patent also notes that the results users select for the same query can change with the time of year: for "Turkey," users normally choose results about the country, but in November they choose results on how to cook a turkey.
Which of these mechanisms does the patent actually claim?
The 9 claims cover only the age classification mechanism. Claims 1, 5 and 9 (method, system and storage medium) describe the classification of related queries as new or old, the new and old counts, the thresholds, and a boost determined from the new count or a penalty determined from the old count. Claims 2, 3, 6 and 7 add the weighting by quality of result statistics. The time trends of extended queries and the popularity change windows are described in the specification and the abstract, but this patent does not claim them.
How does this patent relate to other freshness patents?
This patent reads freshness from queries, while the link velocity patent reads it from links and the document changes patent reads it from content changes. Together they describe 3 independent views of whether a page is current: what people search for, who links to it, and how its text evolves.
What does query freshness change for your SEO?
Query freshness means that the queries that bring visitors to your page tell Google whether it is current. 5 consequences follow:
- Attract the queries of the present. A page clicked from queries with the current year or terms like "latest" or "current" can be classified as new; a page clicked from past years or terms like "previous" can be classified as old.
- Cover the variants users add to recurring queries. For a recurring query like "playoff schedule," the system can favor pages that contain the terms of the most popular extended query of the period, such as "playoff schedule nfl" or "playoff schedule mlb."
- Prepare seasonal content before its window. Pages tied to queries with recurring spikes can be boosted during the spike, so they must be ready and indexed in advance.
- Keep historical content historical. A page about a past edition serves queries that ask for old information, and in some implementations no freshness adjustment is applied to those queries.
- Earn long clicks on current queries. The new and old counts can be weighted by long clicks, so satisfying users on current queries strengthens the new classification.
Our content marketing calendar publishes seasonal pages weeks before their spike, so they are indexed when the recurring queries return.
The patent describes what Google's system can do. It does not confirm how Google adjusts rankings for freshness today.