Semantic distance is a measure of how close 2 terms are in a page according to its structure, such as lists, headings and titles, rather than only the number of words between them. Patent US 7,716,216 B1, "Document ranking based on semantic distance between terms in a document," describes how Google computes it and uses it to rank pages. Google filed it on March 31, 2004 (application 10/813,573), and the patent was granted on May 11, 2010, after a term adjustment of 559 days.
What does patent US 7,716,216 describe?
Patent US 7,716,216 describes a ranking component that detects semantic structures in a page, computes distances between query terms with them, and scores the page's relevance from those distances. Its 2 inventors are Georges R. Harik and Monika H. Henzinger, an inventor of a patent on link history. The patent holds 21 claims and 11 drawing sheets. Its ranking component has 3 parts: a page analyzer that detects lists (including implicit ones), headings and titles; a distance component that computes a distance for each pair of terms; and a relevance component that turns those distances into a relevance score.
The abstract sums it up: the techniques "locate implicitly defined semantic structures in a document, such as, for example, implicitly defined lists in an HTML document," and "the distance values may be used, for example, in the generation of ranking scores that indicate a relevance level of the document to a search query."
Why is word count not enough?
Word count is not enough because the HTML of a page does not follow its visual or logical layout. Search engines consider "the closeness of terms," often measured "by counting the number of words in the document occurring between the search terms." But in web pages with complex formatting, "'closeness' of terms in the underlying HTML file may not correlate with the 'closeness' of the terms when the document is visually displayed."
The patent's example is a list titled "Saturn Facts." Semantically, each row is "probably equally relevant to every other row and to the title of the list": the item "Mass is 95 times that of Earth" should be "considered to be equally distant from the title of the list ('Saturn Facts') as the item 'One Orbit of Sun is 10,759.2 Days,'" whatever their positions.
How does Google detect lists in a page?
Google detects lists by parsing the page into a tree and looking for repeated formatting. Explicit lists use the HTML tags <ul> and <ol>. But "web authors frequently use HTML tags other than <ul> and <ol> to create lists," such as nested tables, <div>, <br> or <p> tags.
The patent detects these implicit lists "by looking for sets of two or more text formatting commands that are repeated." In its example, each item of the Saturn list is a bold word (<b>) followed by a line break (<br>): the repetition of this pair defines the items. Special rules handle the first or last element, for example when the first <br> is missing.
Headings and titles are also detected in the tree, as "text associated with nodes that are above other nodes in the tree structure." The text below a heading or title node "can be considered to belong to" that heading or title.
How does Google measure the distance between 2 terms in a list?
Google measures the distance with 3 rules that adjust the word count, for a list made of a header and items:
- Same item: "if both terms appear in the same list item, the terms are considered close to one another."
- Header and item: if one term is in an item and the other in the header, the pair is "approximately equally distant" to any pair made of the header and another item.
- Different items: "pairs of terms appearing in different list items may be considered to be farther apart" than the pairs of rules 1 and 2.
The patent's illustration: the last word of item A and the first word of item B "may be very close from a word count standpoint," yet "the distance metric may indicate that they are farther apart."
The starting point remains the word count: the distance metrics "may generally be based on how far apart terms are, such as that measured by the number of words," then "further augmented by the concept of semantic closeness." The patent sums up the result: "terms that are in different items of a list, terms that are structurally separated, or terms that are separated by many words are considered far apart."
How do titles and headings change distances?
Titles and headings bring their terms close to the text they govern. "A term in the title of a document may be considered to be close to every other term in document regardless of the word count between the terms." "A term occurring in a heading may be considered to be very close to other terms that are below the heading in the tree structure."
A query term in the title is therefore near every term of the page, and a query term in a heading is near every term of its section.
How does semantic distance change rankings?
Semantic distance changes rankings by favoring pages where the query terms sit close together in the structure. "Documents that include search terms that are relatively close to one another, as given by the distance metrics, may be given higher ranking scores." The patent adds that, "in some applications," other factors can enter the final score, such as "the inverse document frequency (IDF)" to weigh rare words more, and quality measures "based on the interconnecting link-structure of a corpus of documents or those based on a predetermined quality of certain web sites or domains."
The structure analysis "can be performed ahead of time on a corpus of documents" and stored as annotations: at query time, the system then "may simply" look up the pre-calculated annotation information for the document. Structure gives one measure of closeness between terms; word vectors give another, learned from the contexts in which words appear.
The same principle, that headings define what a passage is about, drives the heading context patent for answer passages.
What do the claims of patent US 7,716,216 protect?
The claims protect the list rules, not the title and heading rules. Independent claim 1 requires a list "having a header and a plurality of items associated with the header," and the selection of 1 of 3 rules according to where the 2 terms sit: in different items, in the same item, or in the header and an item. The distance is then computed "using a function based on the selected rule," a function that "differs" for each of the 3 rules, and the result is used to rank the document for a query containing both terms.
The dependent claims add 4 details:
- Word count as a base: claim 6 computes the distance "as a word count, between the first and second terms in the document, augmented by the selected rule."
- Implicit lists: claim 7 identifies them through "repeating occurrences of a set of two or more text formatting commands," and claims 4 and 14 name "paragraph tags, new line tags, bold tags, or table tags."
- Explicit lists too: claims 5 and 21 add the detection of "explicitly defined semantic structures," such as
<ul>and<ol>lists. - Analysis before ranking: claim 11 states that the semantic structure is "identified prior to the ranking."
The rules on titles and headings appear only in the description, as possible implementations.
Where does the ranking component sit in a search engine?
The ranking component comes after retrieval, to sort the documents that already match the query. In the patent's search engine example, the engine first generates "an initial set of documents that match the search query (i.e., documents that contain the terms of the search query)," submits it to the ranking component, then sorts it by the scores it receives.
The patent applies this to "a general search engine," to "a more specialized search engine, such as a news search engine," and to "a corporate document database." A document, in its sense, can be a web page, but also "an e-mail, a blog, a file" or "a news group posting."
What does semantic distance change for your SEO?
Semantic distance means that the structure of your page tells Google which terms belong together. 5 consequences follow:
- Put your main terms in the title. A title term is close to every term of the page.
- Use headings that name the topic of their section. A heading term is close to everything below it, so the heading and its section form one unit.
- Keep related terms in the same list item. Terms in the same item are close; terms split across items are far apart, even when adjacent in the text.
- Give lists a descriptive header. Every item is equally close to the header, so the header connects the list to the query.
- Use clean, consistent markup. Implicit lists are detected from repeated formatting: consistent patterns make your structure readable for machines as well as people.
Our technical SEO work checks that lists, headings and titles use real HTML elements, so the structure the patent reads sits in the markup and not only in the styling.
The patent describes what Google's system can do. It does not confirm that Google measures term distance this way today.