Near-duplicate detection is the identification of pages whose content is the same or almost the same, so that a search engine keeps only one of them. Patent US 6,658,423 B1, "Detecting duplicate and near-duplicate files," describes a fingerprint method that does it at web scale. Google filed it on January 24, 2001, and the patent was granted on December 2, 2003.
What does patent US 6,658,423 describe?
Patent US 6,658,423 describes a method that assigns each document a few fingerprints and declares 2 documents near-duplicates when any one fingerprint matches. Its 2 inventors are William Pugh and Monika H. Henzinger, an inventor of Google's link history patent. The patent holds 38 claims and 18 drawing sheets.
The abstract sums up the method: fingerprints are assigned "by (i) extracting parts from the document, (ii) assigning the extracted parts to one or more of a predetermined number of lists, and (iii) generating a fingerprint from each of the populated lists. Two documents may be considered to be near-duplicates if any one of their fingerprints match."
Where do duplicate pages come from?
Duplicate pages come from 5 sources, according to the patent:
- Mirrors: documents "mirrored" at different sites to reduce delays and network latency.
- Formats: a document can have "plain text and HTML" versions, and more versions as devices multiply.
- Versions: documents with "different information prepended or appended."
- Word replacement: documents "generated from others using consistent word replacement."
- Aggregation: documents that "aggregate or incorporate documents available from another source."
The third source covers information "related to its location on the Web, the date, the date it was last modified, a version, a title, a hierarchical classification path," since "a Web page may be classified under more than one class within the hierarchy of a Web site." The patent's drawings show a concrete case: a search for "aeron chair" returns 5 results pointing to different URLs, on several domains, all showing the same Herman Miller Aeron chair product page with identical product text. The pages "differ only in the date the page was retrieved" ("Monday, July 3" or "Tuesday, July 4"), "the category for the page" ("Home: Personal Care: Aeron Chair" or "Home: Back Care: Chairs: Aeron Chair") "and/or the title." Without detection, all 5 would appear in the results, and "most users would not want to see the others after seeing one."
Removing duplicates serves users, and it lets the search engine "reduce storage requirements" and the "resources needed to process indexes, queries, etc." The patent also notes that duplicates "may indicate plagiarism or copyright infringement."
Why were older methods too expensive?
Older methods were too expensive because they required many shared fingerprints. Earlier techniques fingerprinted elements such as "paragraphs, sentences, words, or shingles (i.e., overlapping stretches of consecutive words)," and treated 2 documents as near-duplicates if they shared more than a set number of fingerprints, at least 2 and "generally much higher." Counting shared fingerprints across "billions of documents" is "quite expensive, computationally and in terms of storage."
The patent adds that discarding fingerprints known to be unique does not help much with those methods, because documents left after that step "may, nonetheless, have no near-duplicate documents."
The new method needs a single matching fingerprint. The patent pairs each fingerprint with its document and sorts the pairs by fingerprint value, so "only documents with matching fingerprints need be" compared, and documents "without any common fingerprints are not checked."
How does the fingerprint method work?
The method works in 3 steps for each document:
- Extract the parts: the document can first be reduced to a "canonical form" without formatting or non-textual components. Parts are then extracted: they "may be sections, paragraphs, sentences, words, or characters," with or without overlap (shingles), and these settings must be "applied consistently across all documents." The patent's example uses single words with no overlap. Stop words can be ignored, and the extraction can skip short documents, "e.g., documents with 50 words or less," because "standard error pages (e.g., informing a user about a dead link, etc.) are typically short, and should not be processed."
- Spread the parts into lists: each part is hashed to decide which of a fixed number of lists it joins. In the patent's example, "the number of lists is set to four," and "each part goes to one and only one list." The hash is "repeatable, deterministic, and not sensitive to state": the word "the" "will always be sent to the same list," whatever the document. For web page text, "three to eight lists" are expected to "yield good results," and good results were obtained "using three lists and four lists."
- Fingerprint each list: each list produces a fingerprint, so a document has as many fingerprints as lists. The fingerprint function must make it "very unlikely that two different lists would produce the same fingerprint," while "two identical lists will always generate the same fingerprint." It can be sensitive to the order of the words in a list, or not.
2 documents are near-duplicates "if two documents have any one fingerprint in common," and could be considered exact duplicates if all their fingerprints are the same. When 2 documents differ by a few words, those words land in only some of the lists: the untouched lists stay identical and still produce the same fingerprint. The number of lists sets the tolerance: as it increases, "the expected number of document differences" needed "before two documents no longer share any common fingerprints increases."
A variant lets each part go into zero, one or several lists, each list having its own independent hash function. If a hash function says "true" for a word with probability p, a list stays unchanged after k different words with probability (1-p)^k, and p and the number of lists can be tuned from there. Because longer documents change more lists, p can be decreased slowly for larger documents so that near-duplicates are still found. Fingerprints that occur in only one document can also be removed beforehand, since they can never match another document.
How does Google use near-duplicate detection?
Google can use near-duplicate detection at 4 moments:
- During crawling: "to speed up the crawling and to save bandwidth by not crawling near-duplicate Web pages or sites, as determined from documents uncovered in a previous crawl." One of 2 near-duplicates is marked as "not to be processed during a subsequent crawl."
- During indexing: "if more than one document are near duplicates, then only one is indexed."
- At query time: documents are grouped in clusters, and if 2 results of the same cluster "match the query equally well (e.g., have the same title and/or snippet)" and "appear in the same group of results (e.g., first page)," only "the one deemed more likely to be relevant (e.g., by virtue of a high PageRank, being more recent, etc.) is returned." The removed result is replaced by the next one: if the fifth of 10 results duplicates the second, the fifth is removed and the eleventh is added. In the Aeron example, results 2 to 5 would be removed and replaced.
- To fix broken links: if a page no longer exists, "a link to a near-duplicate page can be provided."
When duplicates are eliminated outright, the patent gives examples of which one to keep: "the one with best PageRank, with best trust of host, that is the most recent."
How are clusters of near-duplicates built?
Clusters of near-duplicates are built by assuming a transitive property: "if document A is a near-duplicate of document B, which is a near-duplicate of document C, then A is considered to be a near-duplicate of document C," even if A and C would not match directly. Each document is compared with the documents already processed: it joins the cluster of a near-duplicate, 2 clusters are merged when a document links them, and a document with no near-duplicate starts a new cluster. A document "will only belong to a single cluster."
The patent also describes 2 other uses of the fingerprints: shrinking the collection before running any other detection method, and serving as "a pre-filtering step" for a more careful, more expensive technique. In that case, pairs flagged as near-duplicates are checked again, the others are "simply discarded," and the method can be tuned "to err on the side of generating false positive near-duplicate indications."
How does this patent relate to other duplicate-content signals?
This patent removes duplicates from the index and the results, while other patents judge the redundancy of content: Google Panda's guidance asks whether a site has "duplicate, overlapping, or redundant articles," and the information gain patent scores what a page adds beyond pages already seen. The document changes patent uses shingles, overlapping stretches of words that this patent cites among the classic units of earlier fingerprint methods, to measure how much a page changed.
What does near-duplicate detection change for your SEO?
Near-duplicate detection means that when several of your pages say the same thing, Google keeps one and drops the others, and the one it keeps may not be the one you want. 5 consequences follow:
- Give each product one URL. The patent's own example is a product page found at 5 URLs that differ only by category path, crawl date or title: use one canonical URL per product, with categories linking to it.
- Choose your canonical version. The version kept can be the one with the higher PageRank, the more recent one or, when duplicates are eliminated, the one on the most trusted host: tell Google which version you prefer with a canonical tag and consistent internal links.
- Rewrite supplier descriptions. Pages that reuse the same manufacturer text as other stores risk being near-duplicates of them, and then only one version is shown: unique text keeps your page out of their cluster.
- Do not rely on small variations. The patent's near-duplicates differ by a date, a category path and a title, and the method is designed to catch "consistent word replacement": as long as one list stays identical, one matching fingerprint is enough.
- Do not count on very short pages. The patent suggests skipping documents of about 50 words or less, the length of standard error pages: a page that short is treated as not worth comparing.
A technical SEO review from our team lists every URL that serves the same product, then checks that canonical tags and internal links point to the same version.
The patent describes what Google's system can do. It does not confirm which duplicate detection method Google uses today.



