Trust-based ranking is a method that boosts a search result according to the trust placed in the people and organizations who labeled it. Patent US 7,603,350 B1, "Search result ranking based on trust," describes it. Google filed it on May 9, 2006, and the patent was granted on October 13, 2009.
What does patent US 7,603,350 describe?
Patent US 7,603,350 describes a search engine that adjusts retrieval scores with a measure of the trust associated with the entities that labeled each document. Its sole inventor is Ramanathan Guha, also known as R.V. Guha, one of the creators of RSS and of schema.org. The patent has 14 claims and 3 drawing sheets. 3 continuations share its title: applications filed in 2009, 2012 and 2014, granted as US 8,352,467, US 8,818,995 and US 10,268,641.
The abstract describes the full chain: the search engine "determines labels associated with selected documents, and the trust ranks of the entities that provided the labels. The trust ranks are used to determine trust factors for the respective documents. The trust factors are used to adjust information retrieval scores of the documents."
The patent starts from a problem it names directly: on "vertical knowledge sites" (forums, review sites, expert blogs), users know whom to trust, but "when the user returns to a general search engine," that reputation information is lost. The invention brings it into general web search.
What is a label?
A label is a descriptive tag that an entity attaches to web content. The patent defines an entity broadly: "a specific person, group, organization, website, business, institution, government agency or the like." A digital camera expert can label a review of a Canon digital SLR as "Professional review"; an expert can label a group of pages about colon cancer symptoms on a health site as "Symptoms."
Technically, a label lives inside an annotation, written <label, URL_pattern>. The URL pattern can target a single page, a directory or a whole site, and can include wildcards and regular expressions: the annotation "professional review" on www.digitalcameraworld.com/review/ applies the label to every document in that directory. 4 details matter:
- Anyone can label: labels can come from the entity that created the document or from a third party.
- Anchor text can be a label: when an entity simply links to a page, "a crawler extracts the label from the link" and uses the anchor text as the label.
- Labels can carry a score: they "can also include numerical values," a rating or degree of significance.
- Labels are collected two ways: by crawling entities' sites, or through annotation files that entities submit.
The patent describes a query syntax for labels: "label:term." The query "cancer label:symptoms" asks for documents about cancer that have been labeled as relating to "symptoms." When the query contains a label, only the matching labels count. The patent's example: 3 entities label a camera review "professional review" and a fourth labels it "negative review"; for the query label "professional review," "only the first three annotation labels would be deemed matched."
How does Google measure trust between entities?
Google measures trust through explicit and implicit relationships between entities. The patent lists several sources:
- A "trust button" on an entity's site, which records that the visitor trusts that entity.
- Trust lists of entities someone trusts, and vanity lists of entities who trust the site owner.
- Links from a user's web page to the pages of trusted entities.
- Web visitation patterns: a user who visits an entity's page "with a certain frequency" is inferred to trust it.
- Email contact lists and instant messaging chat lists: entities in them can be recorded as trusted.
The patent stores each relationship as a tuple <entity1, entity2, trust_value>, for example <3365, 1230, 1>. In one embodiment, the trust value is 0 or 1.
Trust can be specific to a topic: <entity1, entity2, topic, trust_value>. A user may trust an entity "with respect to politics and economics, but not with respect to sports and entertainment." Trust can also be inferred transitively: if a user trusts a second entity and that entity trusts a first one, the system can assign a trust value between the user and the first entity.
Trust is not permanent. The system can strengthen or weaken a relationship, and can make it "decay over time if the trust relationship is not affirmed by the user," for example by a new click on the trust button. Users can also edit their trust relationships, and crawling periodically updates the trust data.
How is a trust rank computed?
A trust rank is computed from the whole network of trust relationships, an eigenvector calculation comparable to PageRank on the web of links. The system stores trust in a square matrix M, where "each matrix value M_ij stores a value indicative of entity i's trust of entity j." The trust rank of entity i is the i-th component of the eigenvector associated with the eigenvalue 1.
In such a calculation, an entity trusted by many entities that are themselves trusted receives a high trust rank. For topic-specific trust, "a separate trust matrix M is constructed for each topic t," so an entity can rank high in one field and low in another. Topic matrices can also be aggregated along topical hierarchies. Trust ranks are computed in a separate process, ahead of queries, and recomputed when crawling updates the trust information.
How does trust change rankings?
Trust changes rankings through a trust factor applied to the base retrieval score of a document, for example by multiplication. The process of FIG. 4 runs in 7 steps:
- Receive the query: at least one query term, optionally one or more labels.
- Retrieve the documents relevant to the query terms, with any retrieval model (the patent cites PageRank as one).
- Determine the labels of each document, by matching its URL against the URL patterns of the annotations.
- Retrieve the trust ranks of the entities whose labels match the query labels.
- Aggregate the trust ranks per label, then combine the aggregated values into a trust factor.
- Apply the trust factor to the base score, "for example by multiplying with the trust factor."
- Rerank the results by the trust-adjusted scores.
The patent describes several ways to aggregate trust ranks, and says other ways are possible:
- Linear: a fixed weight (for example 1) applied to each trust rank before summing.
- Asymptotic: for example, the sum of the logarithms of the trust ranks, so each extra entity adds less.
- Decaying weight: a weight that decays as the number of instances of the same label grows; for example, trust ranks can be ordered by the age of the annotation, with a decreasing (or alternatively increasing) weight for the oldest one.
- Sigmoid: a sigmoid weighting function.
The patent's example is a review of a Casio digital camera, for a query combining the terms "digital camera" with the label "professional review." The review carries 6 annotations: "Professional Review" by Phil Photo (trust rank 8), Earl Expert (6) and Chris Click (7); "Digital SLR" by Phil Photo (8) and Eddy Shooter (2); "Best buy" by Betsy Buyer (3). Linear aggregation gives 21 for "Professional Review," 10 for "Digital SLR" and 3 for "Best buy." Only the label that matches the query feeds the trust factor, so the 21 counts and the other labels add nothing. The results page shows the name of the entity who provided the matching label next to each result.
Does trust ranking work when the query has no label?
Yes: trust ranking also applies to ordinary queries without labels. The patent states that the system can "receive queries that do not include labels, and still provide trust based ranking." In that case, the trust ranks of the entities behind the annotations applicable to a document are retrieved, aggregated and applied to its base score, and the results are reranked accordingly.
The claims are narrower than the description. Each of the 14 claims requires a query containing a "query label term," described as a "categorical identifier." The claimed method raises the relevance score of a result according to the trust rank of the entity that attached the matching label, ranks the results, and annotates each result with the name of that entity. Claims 3, 5, 9 and 11 add the aggregation of 2 trust ranks into a trust factor.
Is this patent the same as TrustRank?
No: this patent is not TrustRank. TrustRank is an algorithm published in 2004 by Zoltán Gyöngyi, Hector Garcia-Molina and Jan Pedersen, researchers at Stanford University and Yahoo, to separate good pages from spam by propagating trust from a set of seed pages through links.
The 2 ideas differ on their source of trust:
| Point | TrustRank (Stanford and Yahoo, 2004) | Trust patent (Google, 2006) |
|---|---|---|
| Who is trusted | Seed pages | People and organizations |
| How trust spreads | Through links between pages | Through trust declared between entities |
| What it scores | Every page reachable by links | Pages labeled by trusted entities |
| Main goal | Fight web spam | Rank results by expert endorsement |
The link-based approach closest to TrustRank at Google is the trusted seeds model, which measures the link distance between a page and trusted seed pages.
What does the trust patent change for your SEO?
The trust patent means that a page gains ranking strength when trusted entities label it with terms that describe it. 6 consequences follow:
- Get endorsed by experts in your field. The trust factor depends on the trust rank of the entities who label your page, not only on how many of them do it.
- Build trust topic by topic. In the topic-specific version, each topic has its own trust matrix: authority in one field does not automatically carry over to another.
- Earn the right label, not any mention. When the query contains a label, only matching labels count. Being labeled "professional review" matters for that query; a label from an unrelated category adds nothing.
- Watch your anchor text from trusted sites. A link from an entity's site can become an annotation, with the anchor text as the label: a descriptive anchor from a trusted expert works as an endorsement on that topic.
- Keep trust relationships alive. Trust between entities can decay over time if it is not reaffirmed, so a reputation earned once is not guaranteed forever.
- Make your authors identifiable. Trust ranks belong to entities, and the results show the name of the entity behind the matching label: a named expert with a public track record can accumulate trust, which an anonymous source cannot do.
The same principle of trust earned from independent sources runs through Google Panda, which counts only links from unrelated sites.
The patent describes what Google's system can do. It does not confirm that Google uses labels or trust ranks in its search results today.