Google's generative summaries are answers written by a large language model (LLM) from the content of selected search result documents, with each verifiable portion linked to a page that supports it. Patent US 11,769,017 B1, "Generative summaries for search results," describes the system. Google filed it on March 20, 2023, from a provisional application of December 30, 2022, and the patent was granted on September 26, 2023.
What does patent US 11,769,017 describe?
Patent US 11,769,017 describes a search system that feeds content from search results into an LLM to write a natural language summary of the answer. Its 10 inventors are Matthew K. Gray, John Blitzer, Corinn Herrick, Srinivasan Venkatachary, Jayant Madhavan, Sam Oates, Phiroze Parakh, Aditya Shah, Mahsan Rofouei and Ibrahim Badr. The patent holds 22 claims and 12 drawing sheets, and it names 2 Google models as examples of LLMs: PaLM and LaMDA.
The query can be typed, spoken, based on an image, multimodal (the patent's example: a photo of an avocado with the voice input "is this healthy"), or implied: generated automatically from context or profile data, such as "patent news" for a user interested in patents. The summary can be displayed or read aloud.
The patent never uses the names "Search Generative Experience" or "AI Overviews." Google presented its Search Generative Experience at Google I/O on May 10, 2023, less than 2 months after the filing, and announced the launch of AI Overviews in the United States on May 14, 2024. The patent describes the mechanism behind this kind of answer.
Which problems does the patent solve?
The patent solves 3 weaknesses of an LLM that answers from its training data alone:
- Inaccurate answers: the LLM may not be trained on fresh data, may be trained on inaccurate data, or may produce "a so-called hallucination." The patent's example is the query "how to change DNS settings on Acme router": the answer "the default IP address is 192.168.1.1" can be wrong if the model learned it from stale data and the default address has since changed.
- The same answer for every user: an LLM returns the same output to a non-tech-savvy user and to a tech-savvy user who submit the same query.
- Over-specified or under-specified answers: the answer gives the expert more than needed, or gives the beginner too little to finish the task.
The fix is to process additional content with the query, or even without it: content taken from search result documents, which can be "more up to date than training data on which the LLM has been trained." The patent states a limit, though: the summary can still include a portion "not derivable directly from the corresponding content processed using the LLM," generated from the model's prior training.
Which documents does Google feed to the LLM?
Google feeds the LLM a set of search result documents drawn from 4 sources:
- Query-responsive documents: the results for the query itself. The system can, for instance, keep the top N results (the patent cites 2 or 3 as examples).
- Related-query-responsive documents: results for related queries, such as queries "often issued, among a population of users, in close temporal proximity to the query."
- Recent-query-responsive documents: results for queries recently submitted from the same device or account, favored when they were issued close in time and overlap in topic or entities with the query.
- Implied-query-responsive documents: results for queries generated automatically from context or profile data.
The 3 extra sources are optional and selective. In the patent's example, related-query results are added "only when" the query's own results "are of low quality and/or not diverse relative to one another, and the magnitude of the correlation satisfies a threshold." That correlation can be based on how often both queries are issued by the same device or account close together in time. Recent and implied queries follow the same logic, with an overlap threshold. The goal is to spend extra computing resources only when they are likely to produce a more accurate summary.
The prompt can include the query, "In the context of <query>, summarize <Content A>, <Content B>, <Content C>, and <Content D>," or omit it: "Summarize <Content A>, <Content B>, <Content C>, and <Content D>." In the example, A and B come from query results, C from a related query and D from a recent query.
Which measures decide that a page becomes a source?
3 families of measures can decide which search result documents enter the set. The patent lists them with examples.
| Family | Measures named in the patent |
|---|---|
| Query-dependent | Positional ranking for the query, selection rate for the query, locality (origin of the query vs location of the document), language match |
| Query-independent | Selection rate across many queries, trustworthiness (based on the author, the domain or inbound links), overall popularity, freshness |
| User-dependent | Relation to the attributes of the user profile, to recent queries and to recent non-query interactions of the user |
The freshness measure "reflects recency of creation or updating" of the document. The trustworthiness measure can be "generated based on an author thereof, a domain thereof, and/or inbound link(s) thereto." For user-dependent measures, the patent's example is a "movie buff": a page about a movie is more likely to enter that user's set.
These measures make the set vary from one submission of a query to another. The query "history of Louisville" sent from Louisville, Kentucky, and from Louisville, Colorado, can receive 2 distinct sets of documents, and therefore 2 different summaries.
Which part of a page does the LLM read?
The LLM reads all of a selected page or a subset of it, chosen for its correlation with the query. The content can be text, images (through generated captions, text detected in them or descriptions of detected objects) or videos (through transcriptions). The patent describes a snippet selected because it contains query terms, terms similar to the query terms, or because "a word embedding of the snippet" is "within a threshold distance of a word embedding of the query." Word embeddings of this kind descend from the Word2vec patent, which places words in a space of meaning.
The system can summarize that content before passing it to the LLM, for example when the content does not fit the memory constraints of the LLM. It can also add a source identifier to each piece: a token at the beginning or the end of the content, descriptive or generic such as S1, S2, S3.
How does Google link a sentence to its source?
Google links a portion of the summary to a source when a document verifies that portion. The patent calls the process "linkifying." A portion can be a sentence, a semantically coherent part of a sentence, or a span of N characters or N words. The patent describes 2 ways to decide that a document verifies it:
- The LLM output itself: a source identifier (S1, S2...) placed before or after a portion indicates that the matching document verifies it.
- An embedding comparison, in 4 steps:
- Encode the portion of the summary into a content embedding.
- Encode a portion of the candidate document into a document content embedding.
- Measure the distance between the 2 embeddings.
- Add the link when "the distance measure satisfies a threshold."
The candidate document comes either from the documents used to write the summary or from a new search run on the portion itself, keeping the top result or one of the top N results. The linked document can therefore be one the LLM never read. The system stops after a verifying document or after N candidates, but a portion can also receive several links, each shown as its own icon. A portion with no verifying document stays without a link.
The user then reaches the source with "a single tap, single click, or other single input," and the link can be an anchor link to the passage that verifies the portion. In the patent's figures, the 3 sources cited in the summary have different URLs from the 3 classic results shown below it.
How confident is Google in a summary?
Google measures confidence per portion of the summary and for the summary as a whole, from the LLM's own confidence and from the documents that verify it. For a portion, confidence can depend on the number and the trustworthiness of the verifying documents. The patent gives 2 comparisons: confidence is greater "when four SRDs verify that portion than when only one SRD verifies that portion," and greater "when four highly trustworthy SRDs verify that portion than when four less trustworthy SRDs verify that portion."
Confidence decides whether and how the summary is shown:
- Above an upper threshold: the summary can be shown without any classic results at first; they appear after a delay or on request.
- Between the 2 thresholds: the summary is shown with the classic results.
- Below the lower threshold: the summary is suppressed and only the classic results are shown.
The interface can also annotate confidence, with a "high confidence" or "low confidence" label for the whole summary, or a color per portion (green, orange, red in the example).
Before any of this, a step selects "none, one, or multiple" generative models for the query, using classifiers or rules: an informational LLM, a creative LLM for poems or essays, a text-to-image diffusion model, or a smaller or larger LLM. The system can select no model when, for example, the "search results for the query are high quality." A query can therefore receive no generated summary at all.
How does the summary adapt to each user?
The summary adapts to what the user already knows and to the results the user has opened. The system compares the profile with the query and its results: the user is considered familiar with some content after interacting with documents that contain it, or when the profile directly indicates it. The prompt then becomes "assuming the user is familiar with [description of the certain content] answer [query]."
After the user interacts with a result, the system generates a revised summary, with additional content such as "assume the user already knows X." The revised summary can omit what that result covered, or on the contrary discuss it in more depth. An interaction can be a click with a threshold dwell time, but it does not require a click: the patent counts "pausing of scrolling over the search result, expanding of the search result, highlighting of the search result."
The revised summary can replace the initial one when the user comes back to the results page, or appear when the user submits the same or a similar query again, "minutes, hours, or days later."
What do the 22 claims of the patent protect?
The claims protect 2 independent methods:
- Claim 1: the summary revised after an interaction with a search result document, using an input that reflects that interaction (claims 2 to 12 add viewing for a threshold duration, replacement of the initial summary, a prompt reflecting familiarity, the selection of query and related-query results).
- Claim 13: the selection of a set of search result documents, the summary generated from their content, and a confidence annotation on a portion of the summary (claims 14 to 22 add selection by query-dependent, query-independent and user-dependent measures, related queries above a correlation threshold, links on verified portions, generation without the query, and implied queries).
The other mechanisms (the 4 sources, embedding verification, model selection) are described in the patent but are not all claimed as such.
Which other variants does the patent describe?
- Output formats: the user can ask for a "list format," a "graph format," a "top 5" or a summary "in the style of" someone, which selects an LLM fine-tuned for that format or adapts the prompt.
- Several LLMs in parallel: each writes a candidate summary, and the system keeps one, for example the one most similar to the other candidates.
- Several LLMs in series: a first LLM selects passages, a second summarizes each passage, a third writes the overall summary.
What does the generative summaries patent change for your GEO?
The patent means that a page gets cited in a generative summary when search selects it, when its passages verify precise statements, and when trustworthy sources say the same thing. 6 consequences follow:
- Rank in the classic results first. The LLM reads search result documents selected with measures such as positional ranking and selection rate, and the verification search also keeps the top results for the sentence. A page that ranks for nothing relevant has no path to becoming a source.
- Write passages that verify one statement each. With the embedding method, a link appears only when the embedding of a passage sits close to the embedding of a portion of the summary. Our reading: short, factual, self-contained passages are the easiest to match.
- Build trustworthiness signals. The patent names the author, the domain and inbound links as the basis of the trustworthiness measure, which also weighs on confidence.
- Update your pages. The freshness measure reflects the recency of creation or update, and fresher content than the LLM's training data is one of the reasons the patent gives for reading web pages at all.
- Cover related and follow-up queries. Results for related, recent and implied queries can enter the set when the query's own results are weak or not diverse, and revised summaries move on to what the user has not read yet.
- Be corroborated. Confidence grows when several trustworthy documents verify the same portion, and a low confidence can remove the summary entirely: state facts that independent, trustworthy sources confirm.
Our GEO service checks these 6 points page by page, starting with whether the page ranks in the classic results for the queries a summary draws on.
The click data that feeds selection rate is the core of Navboost, and trustworthiness echoes Google's patent on trust, which ranks pages by the entities that vouch for them.
The patent describes what Google's system can do. It does not confirm which measures AI Overviews use today, or their weights.