A contextual synonym is a word or phrase that can replace another in a query only when the surrounding words allow it. Patent US 7,636,714 B1, "Determining query term synonyms within query context," describes how Google learns such synonyms from the way users rewrite their own queries. Google filed it on March 31, 2005, and the patent was granted on December 22, 2009.
What does patent US 7,636,714 describe?
Patent US 7,636,714 describes a method that mines query logs for pairs of queries that differ by one phrase, then tests whether the swapped phrases are synonyms in a given context. Its 2 inventors are John Lamping and Steven Baker; Steven Baker also co-signed the heading context patent for answer passages. The patent holds 19 claims and 4 drawing sheets, and was assigned to Google Inc.
The abstract sets out the method: queries are "sorted by user identity and session"; for each query, "a plurality of pseudo-queries is determined," each "derived from a user query by replacing a phrase of the user query with a token"; candidate synonyms are terms "used within a user query in place of the phrase"; their strength is evaluated; validated synonyms "may be either suggested to the user or automatically added to user search strings."
Where do the synonyms come from?
The synonyms come from users who reformulate their own searches. Each query is stored "with a user identifier," "a timestamp, and a list of some number of the search results (e.g., a list of the top ten document IDs from the search)." All queries received over a period of time, "such as a week," are sorted "by user ID (e.g., by cookie), and then by time." A session is defined as the queries from one device within a given interval, "for example one hour." The system then looks for pairs of queries that vary by a single phrase.
The patent's example of such a pair: [free loops for flash movie] followed by [free music for flash movie]. Queries with too little context can be dropped: queries used in the analysis "may be required to have at least three terms." In the patent's session example, the query [gm cars] is eliminated for this reason.
What is a pseudo-query?
A pseudo-query is a query in which one phrase is replaced by a token that acts as a variable. The patent uses ":" as the token. Every sequence of one or more consecutive terms is replaced in turn, "while leaving at least two words in the pseudo-queries." From [gm used car prices], the system forms 4 pseudo-queries: [: used car prices], [gm : car prices], [gm used : prices] and [gm used car :]. Every query that fits this pattern, such as [general motors used car prices], becomes a candidate pair, and the phrase that fills the token, "general motors," becomes a candidate synonym of "gm."
The patent's worked example compares the top ten results of the queries. In the context of [gm used car prices], "general motors" shares 5 result URLs with the original query, while "ford" shares none. In the context of [nutrition of gm food], it is "genetically modified" that shares 6 result URLs, while "macdonalds" shares none. The number of shared results separates the real synonym from the false one, and the right synonym of "gm" changes with the surrounding words.
Before any context analysis, each phrase keeps its ten most significant candidates, ranked by a hierarchy of tests: first the number of related queries within sessions, then the number of results in common, then the frequency within the queries of the pseudo-query. The patent counts distinct queries rather than repetitions: "it is more meaningful to examine many distinct queries than to simply count multiple occurrences of a given query."
How does Google test a candidate synonym?
Google tests a candidate synonym with evidence from sessions and from search results. The patent lists 3 kinds of evidence and tests:
- Co-occurrence in sessions: "the frequency with which both queries in the pair are asked by the same user within a short time interval." In one embodiment, "the interval is an hour and the probability is 0.1% or greater."
- Probability of the swap: for every query containing phrase A, the query with phrase B substituted "has a moderately high probability of occurrence in the stored data." In one embodiment, the required probability is 1%.
- Shared results: the 2 queries have "a minimum probability of having a number of the top results in common," for example a probability of 60 to 70% with 1 to 3 results in common.
Threshold conditions can apply first, for example that "for at least 65% of the original-altered query pairs, there is at least one search result in common," and that the altered query follows the original "within five sequential queries" at a frequency of "at least 1 in 2000." The 65% parameter "is empirically derived, and other thresholds can be used as well, depending on the corpus of documents." The patent adds that statistics "should be gathered over at least 1000 queries including the phrase" to have confidence in them.
How does Google score the strength of a synonym?
Google scores a synonym by combining 4 tests into a single confidence value called evidence. Each test compares a measured ratio to a target value through a Scale function, which returns 0 when the ratio equals the target and approaches 1 as the ratio grows. The 4 tests are:
- frequently_alterable: for queries with the phrase, the altered query occurs often enough in the logs, "preferable more than 1%."
- frequently_much_in_common: original and altered queries share enough results; "Preferably, at least 60% of altered queries have at least 3 search results in common with the original user query."
- frequently_altered: users occasionally try the substitution; "Preferably, for every 2000 user queries, there is a corresponding altered query within the same session."
- high_altering_ratio: users do not preferentially substitute in the opposite direction, which "would suggest that the original phrase is much better than the candidate synonym."
The formula weights the tests: soft_and = frequently_alterable + 2 × frequently_much_in_common + 0.5 × frequently_altered + high_altering_ratio, then evidence = 1.0 - exp(-soft_and / 1.5). Shared results carry the heaviest weight. "A value approaching 1.0 indicates very high confidence, while a value of 0.6 reflects good confidence," and for many applications a candidate can be validated "if the value of evidence is greater than 0.6." The threshold depends on the application.
Why does context matter?
Context matters because a word can be a synonym in one setting and not in another. The patent's example uses the pair [killer whale free photos] and [killer whale download photos]. The swap "free" to "download" appears in 4 contexts: after "whale" (whale :), before "photos" (: photos), between the 2 (whale : photos), and in general (:). These contexts are adjacent words; for synonyms in long queries, a later Google patent checks the candidate against words far apart in the query.
Statistics are gathered for each context in which a phrase occurs frequently, for example "the 10,000 contexts for which the most queries exist," keeping for each context "only the 20 most common candidate synonyms." The tests can then run for each context separately. The patent's conclusion: "'download' is not generally (i.e., in the general context) a good synonym for 'free,' is a good synonym in the context (: photos), and is not a good synonym in the context (: press)." The patent calls the context (: photos) "an exception to the general rule." Free photos and downloadable photos serve the same need; free press and download press do not.
This is why the patent rejects fixed lists. A thesaurus is "expensive to construct," "generally restricted to one language," and misses contextual synonyms: "music" is not listed as a synonym for "loop" in standard thesauruses, yet it is one in [free loops for flash movie]. Clustering "related words" fails the other way: "sail" and "wind" occur together in many documents, "but they are not synonymous."
How does Google use the synonyms?
Google uses validated synonyms to suggest alternative queries or to expand the user's query automatically. The patent's flowchart shows the search: receive a query, identify results, determine candidate synonyms for a phrase from a predetermined list, derive an altered query, and identify results for the altered query. The patent describes 3 levels of use:
- Suggestion (the "conservative approach"): alternative queries containing the synonym are shown with the results of the original query, each linked to its own results.
- Automatic expansion (the "more aggressive approach"): the phrase is replaced by a disjunction, for example "gm" would be replaced by "gm" OR "general motors." A query with "gm" can then retrieve pages that only say "general motors." This disjunction is the core of claim 1.
- Score adjustment: the synonym can be used "solely to modify the score associated with the retrieved documents"; claim 11 modifies the ranking of results "based on whether the search results include the replacement term."
If the evidence for a synonym is relatively weak, the synonym can be used "as suggestive rather than equivalent."
The idea that relevance depends on the words around a term runs through phrase-based indexing, where related phrases predict each other.
What does contextual synonym learning change for your SEO?
Contextual synonym learning means that Google matches your page to the words people mean, not only the words they type, but only where the context allows it. 5 consequences follow:
- Stop creating a page per synonym. If users treat 2 phrases as interchangeable in a context, Google can expand one query to the other: one strong page serves both.
- Use the wording of your audience. Synonyms come from how real users reformulate: their vocabulary, including abbreviations like "gm," is the one Google learns. Their queries carry time signals: Google reads the dates and terms in them to decide when fresh results matter.
- Respect context-specific meanings. A synonym valid in one context can be invalid in another: do not assume that a word swap works across all your topics.
- Check shared results before targeting a variant. The patent validates synonyms by shared top results: if 2 queries return the same pages, they are one target; if not, they are 2.
- Write for intent with natural variation. Covering a topic with the terms people use around it lets query expansion work in your favor.
The patent describes what Google's system can do. It does not confirm how Google learns synonyms today.