Link spam detection by derivative is a method that spots pages whose PageRank is artificially inflated, by measuring how their importance changes when the weight given to links changes. Patent US 7,953,763 B2, "Method for detecting link spam in hyperlinked databases," describes it. The original application was filed on August 18, 2004, from a provisional application of August 18, 2003. This continuation (application 12/410,381) was filed on March 24, 2009 and granted on May 31, 2011.
What does patent US 7,953,763 describe?
Patent US 7,953,763 describes an inflation detector that computes, for each page, a quantity corresponding to the mathematical derivative of its importance (its PageRank, in the main embodiment) and flags pages whose value is abnormal. Its 3 inventors are Sepandar D. Kamvar, Taher H. Haveliwala and Glen M. Jeh, researchers known for their work on PageRank computation. Sepandar Kamvar also co-signed the author reputation patent. The patent holds 25 claims and 6 drawing sheets, and it continues patent US 7,509,344. In the search engine described, the inflation detector examines the link maps and uses the PageRanks, and it "may alter the PageRanks 132 or the link maps 128 as a result of detecting inflated nodes."
The patent addresses a known weakness of link-based ranking: "the link structure surrounding a node can be deliberately modified to artificially inflate the rank of the node."
Which kinds of link spam does the patent target?
The patent targets 2 kinds of link spam:
- Link farms: "a set of nodes where a large number of nodes point to a single node in order to give the false impression that the single node is important." The patent's example is the home page of a commercial site, boosted by "many dummy web documents" that all link to it.
- Clique attacks, or web rings: "a set of nodes that point predominantly to one another to give the false appearance of authority or importance." Its example is a ring of 4 pages with many mutual links.
Why does a simple link count fail?
A simple link count fails because real authorities look like link farms. A link farm is "one central node with many other nodes linking to it," but "this same structure, however, naturally occurs in linked databases whenever a very important node is linked to by many other nodes." The patent's example: "the web site Yahoo.com has many links pointing to it, but it is not a link farm."
The difference lies in who links. For a real authority, "the nodes linking to the central node tend to have some links with relatively high rank," while in a link farm, "the nodes linking to the central node all tend to have relatively low rank." Checking every node for this pattern is "computationally prohibitive" at web scale, so the patent needs a quantity that captures it directly.
How does the derivative detect link spam?
The derivative detects link spam by measuring how a page's importance reacts to the link coupling factor, which plays the role of PageRank's damping factor: in the matrix A(c) = [cP + (1 − c)E]ᵀ, c weighs the real links (P) against random jumps (E). The coupling factor ranges from 0 to 1:
- At 0, nodes are "completely decoupled": users jump at random and "all nodes are assigned an equal rank."
- At 1, nodes are "completely coupled": there are no random jumps, and ranks depend entirely on links.
The system computes the derivative of each page's importance with respect to this factor, then, in some embodiments, normalizes it by the page's importance. The importance used for normalization "is not necessarily the same importance of the node for which the derivative is taken": it can also come from a count of in-links, a principal eigenvector or a singular value decomposition of the link matrix. The 2 kinds of spam give opposite signatures, measured "in comparison with the same quantity for other nodes in the graph":
| Structure | Normalized derivative | Why |
|---|---|---|
| Link farm | Large and negative | The many linkers have very low importance, so the central page loses fast as links weigh more |
| Real authority | Moderate | High-ranked and low-ranked linkers cancel each other out |
| Web ring | Large and positive | Importance circulates inside the ring and is not dissipated to outside pages |
| Natural cluster | Moderate | It has more links to outside pages, which dissipates the mutual reinforcement |
The patent computes the derivative from the PageRank equations, x'(c) = (I − cPᵀ)⁻¹(P − E)ᵀx(c), where x(c) is the importance vector. Because the matrix "tends to be very large and sparse," a factorization is "prohibitively expensive," so one embodiment solves the linear system with "a Jacobi relaxation technique," a simple iterative method whose convergence rate equals c, which is very fast for values of c up to 0.98. Gauss-Seidel techniques are named as an alternative.
Does the patent measure the derivative at a single value of the factor?
No, not necessarily. Because "the derivative function may not be uniform over all values," some embodiments compute several derivatives for each page at different values of c, then combine them, for example by averaging, to capture the change "over a wider range." When the derivative is normalized by the page's own rank, the average over an interval [a, b] has a closed form: [log x(b) − log x(a)]/(b − a). In practice, the spam likelihood can be computed directly from the page's importance evaluated at 2 values of c, as the abstract states.
Can the derivative also rank pages?
Yes. The patent states that "irrespective of link spam considerations, the derivative value may be used to assign a rank to a node," for example to sort search results. Claims 20 to 25 cover this use: in response to a search query, the system orders the pages "in accordance with the respective quantities" and returns the results in that order. The method is also not limited to the web: the patent lists journal articles, patents citing other patents, newsgroup postings, email messages, social networks and peer-to-peer networks.
What happens to pages flagged as link spam?
Flagged pages become candidates for counter-measures. The patent describes several selection rules: a predetermined percentage of pages with the lowest normalized derivative for link farms, the highest for web rings, or the largest magnitude to catch both, or a threshold on these values. "A human or supplementary algorithm may be used to examine the possible link spam nodes to make a final determination."
The counter-measures are explicit: "the node is eliminated from the graph," or its importance is reduced by "a predetermined penalty or a calculated amount, e.g., an amount proportional to the magnitude of the normalized derivative value." The adjustment can apply to a ranking computed by other techniques than link-based ranking, or by a combination of both. In the main independent claims (1, 12 and 19), performing "a remedial action" on the identified pages is part of the method itself.
How does this patent relate to other link spam patents?
This patent detects spam from the shape of the link graph. The link context patent detects it from the words around links, the link velocity patent from the timing of links, and seed-based PageRank makes farms useless by measuring distance from trusted pages.
What does link spam detection change for your SEO?
Link spam detection means that links from many weak pages and closed circles of mutual links leave a mathematical signature that separates them from real authority. 5 consequences follow:
- Never build link farms. Many low-importance pages pointing to one page produce a sharply negative signature.
- Avoid link rings and reciprocal networks. Groups of sites that mostly link to each other produce a sharply positive signature.
- Earn some links from strong pages. What distinguishes a real authority is the presence of high-ranked linkers among the many small ones.
- Keep your link profile open to the outside. Natural sites link out to other sites, which dissipates the mutual reinforcement that defines a ring.
- Penalties are possible, not just neutralization. Flagged pages can be removed from the graph or have their importance reduced in proportion to the anomaly.
A backlink audit from our team looks for both signatures: many weak pages pointing to one URL, and groups of sites that mostly link to each other.
The patent describes what Google's system can do. It does not confirm that Google uses this derivative to detect link spam today.



