An information gain score is a measure of the new information a document adds beyond the documents a user has already seen on the same topic. Patent US 11,354,342 B2, "Contextual estimation of link information gain," describes how Google computes it and uses it to rank and present documents. Google filed the international application (PCT/US2018/056483) on October 18, 2018, the US application was published as US 2020/0349181 A1 on November 5, 2020, and the patent was granted on June 7, 2022.
What does patent US 11,354,342 describe?
Patent US 11,354,342 describes a system that scores unseen documents by the information they would add for a given user, then selects or re-ranks them with that score. Its 2 inventors are Victor Carbune and Pedro Gonnet Anders, both based in Switzerland. The patent holds 16 claims and 6 drawing sheets.
The abstract gives the definition: an information gain score "is indicative of additional information that is included in the document beyond information contained in documents that were previously viewed by the user." In the claims, the score is "based on a quantity of new information included in the given new document" that "differs from the information extracted from each document of the first set." The patent applies the score to 2 surfaces: a search results page and an automated assistant that presents information, often read aloud through text-to-speech. Claim 1 covers the assistant case; claims 5 and 10 cover documents the user accessed, including through a search results interface.
Which problem does information gain solve?
Information gain solves the redundancy of documents that share a topic. The patent's example is a search about a computer problem: the user receives "multiple documents that include a similar listing of solutions, remedial steps, resources, etc." 2 documents can both be relevant to the query, yet the user "may have less interest in viewing a second document after already viewing the same or similar information in a first document or set of documents."
Relevance asks whether a document matches the query. Information gain asks whether the document still teaches the user something.
How does Google compute an information gain score?
Google computes the score in 5 steps, shown in the patent's main flowchart (FIG. 5):
- Identify a first set of documents that share a topic provided by the user and that the user has already viewed (or listened to).
- Identify a second set of documents on the same topic that the user has not viewed.
- Determine an information gain score for each document of the second set, in some implementations with a machine learning model.
- Select a new document based on the scores.
- Present information from the selected document to the user.
The model receives data from both sets: the documents themselves or a semantic representation of them, such as "an embedding, a feature vector, a bag-of-words representation, a histogram generated from words/phrases in the document." In one setup, n inputs are reserved for semantic features of already-presented documents and m inputs for the new document, for a dimension of n + m. Unused inputs can be left blank or filled with zeros, so they have little or no impact on the output.
The output can be "a quantitative score between 0.00 and 1.00," where 0.00 means that no information gain is expected given the documents already viewed, and 1.00 means that the new document contains only information the user has not seen (the patent calls it "total information gain"). The patent also allows other forms, such as "a value along a range of numbers."
Is the score the same for every user?
No. The score is computed for one user, against the documents that user has already been presented. A user document database stores data about the documents presented to the user, such as entities, a location identifier or a semantic vector of the text. When none of the documents has been viewed yet, the patent says an information gain score "may not be generated" for any of them, or an arbitrary equal score may be assigned. This per-user database resembles the user profile that Google builds for personalized search from terms, categories and links.
The scope can be a single session: in some implementations, "each time a topic is provided by the user" starts "an independent set of viewed documents that is not persisted through multiple topic submissions." In others, the references are stored after a search session ends.
Information gain does not replace relevance. The patent describes documents that are first scored by the search engine, then "the scores may be adjusted based on information gain scores." The ranked list can be based "on the information gain scores and/or one or more other scores."
How is the model trained?
The model can be trained on pairs of documents labeled with the information gain of the second over the first. Each training pair is applied to the model, the output is compared to the label, and the error trains the model, with techniques such as gradient descent and back propagation. The patent names several model types: neural networks ("feed-forward, convolutional, etc."), "support vector machines, Bayesian classifiers." Claim 15 specifies a neural network.
The labels can come from people, in several ways:
- Individuals read the documents and give "a subjective information gain score" for how much new information they gained from the second document after the first.
- Human curators assign a value to the second document of a pair, through an information gain annotation engine.
- People who simply search in the ordinary course of their lives can be asked, for example "via a web browser plugin," questions such as "[Was] this document/information helpful in view of what you've already read?" or "Was this document redundant?"
- A user in a search session can see a pop-up asking to rate the information gained from a viewed document.
In these setups, the labels encode a human judgment of what is new and what is redundant.
The semantic vectors themselves can come from another trained model: the patent cites an autoencoder (e.g., word2vec) whose encoder part generates the semantic representation of each document.
How does information gain change the results page?
Information gain changes the results page as the user reads. The patent's example uses 4 documents for the query "Help me fix my computer." Document 1 covers common software troubleshooting, Document 2 adds common software application issues, Document 3 covers hardware troubleshooting and software solutions, and Document 4 covers only hardware solutions. Because Document 1 and Document 2 overlap, Document 2 gets a score less indicative of information gain than Documents 3 and 4, and Document 4 gets the highest, since its information is "not included at all" in the others.
When the user opens Document 1 and returns to the results, the remaining documents are re-scored and re-ranked: "Document 4 may now appear higher in the list than Document 2." Each new document viewed moves from the second set to the first, and the scores are computed again.
The patent goes further. A document found to include "identical information" as another document "may be removed from the list (or at least demoted substantially)" when the user navigates back, which may indicate "a determined zero information gain." The viewed document itself can be excluded from the refreshed list (claims 8 and 13). Pages that are near copies of each other face a comparable filter at indexing time, where Google fingerprints pages to keep only one version.
How does an assistant use information gain?
An automated assistant applies the same logic to skip what the user has already heard. The patent's example is a spoken question about a computer error message. A first document gives 2 information elements, 2 possible causes of the error, both read aloud through text-to-speech. After a follow-up such as "what else could be the cause of this error message?", a second document contains the second cause and a third one. The assistant determines that "the second information element has already been conveyed to the user" and reads only the third.
The patent explains why this matters for audio: listening takes longer than reading and the user cannot easily scan or skip, so removing redundant information shortens the output and can reduce the number of user interruptions and extra dialog turns. On a results page, the stated benefit is a shorter or more efficient query session, with "fewer input interactions."
Is information gain the same as Google's helpful content guidance?
No, not literally: the term "information gain" does not appear in Google's guide to creating helpful, reliable, people-first content, but both point in the same direction. The guide asks whether content provides "original information, reporting, research, or analysis" and "insightful analysis or interesting information that is beyond the obvious." The patent's score expresses a related idea, computed against what a given user has already seen.
The patent describes what Google can compute, not a confirmed ranking factor. The redundancy question echoes the Panda patent, whose 2011 guidance asked whether a site has "duplicate, overlapping, or redundant articles."
What does information gain change for your SEO?
Information gain means that a page earns its place by what it adds to the pages that already rank, not by repeating them. 5 consequences follow:
- Read the top results before writing. List what every ranking page already says: in the patent's logic, that is the information a reader who has seen one of them would not gain again.
- Add information the others do not have. First-hand data, tests, original examples, numbers from your own work and expert analysis are the content no competitor can duplicate.
- Answer the next question. A reader who has seen the first result still has questions. Cover the follow-up questions the top results leave open, as the assistant in the patent does with "what else could be the cause?"
- Do not rewrite the consensus. In the patent's terms, a page that paraphrases the top 3 results would bring little information gain to a reader who has seen one of them, and a page with identical information can be removed or demoted.
- Structure the unique parts. Put the original information in clear, self-contained sections, so that an assistant or a generative summary can extract it as the part that adds value.
Our content marketing starts with step 1: we list what the top 10 results already say, then build each page around the first-hand data and follow-up questions they leave open.
The patent describes what Google's system can do. It does not confirm that Google uses an information gain score in its rankings today.