Information gain

nounalso called information gain score or net-new information

Definition

Information gain is the new information a page adds beyond what other pages on the same topic already say — original data, first-hand experience, a new angle — and it is the leading explanation for why content that rewrites the top results struggles to rank or be cited.

Updated 5 min read8 cited sources

On this page7 sections

Why it matters for founders and small teams

If your article says what the top five results already say, an AI engine has no reason to cite you instead of them, and in a tie the bigger brand wins. Information gain is how a small team competes without a bigger budget: one number, test or first-hand lesson that nobody else has published gives Google and AI engines something they can only get from you.

Is information gain a Google ranking factor?#

Information gain is not a confirmed Google ranking factor: the term comes from a Google patent, granted in 2022, that scores documents by what they add beyond pages a user has already seen — but Google has never said it uses that system, and a patent shows an invention, not a deployment.

The patent is US11354342B2, “Contextual estimation of link information gain,” filed by Google in October 2018 and granted in June 2022. Its first claim describes an automated assistant that has already presented a user with information from some documents on a topic. It scores new documents on the same topic for the “additional information that would be gained by the user” beyond what they’ve already seen, and presents information from the one it selects. The description adds that documents can be re-ranked as the user views more of them.

  • The score is relative to one user. It compares a document with the documents that user has already been shown, not with every page on the web.
  • It’s framed around an assistant. The claims describe a system answering free-form natural-language questions — closer to today’s AI search than to ten blue links.
  • It’s unconfirmed. Google hasn’t said whether, or where, this system runs. Treat claims of an “information gain score” in Google’s rankings as speculation.

What Google does confirm is the principle. Its helpful-content guidance asks whether a page provides “original information, reporting, research, or analysis” and whether it avoids “simply copying or rewriting” its sources, and its ranking systems include original content systems that show original reporting “ahead of those who merely cite it.”

Information gain matters more in AI search because a language model can already write the consensus answer from its training data and a few sources, so the pages worth citing are the ones that supply what it can’t generate: a figure, a test result, a first-hand account or a primary source.

  • Non-commodity content

    Official

    Google’s AI features guide contrasts commodity content like “7 Tips for First-Time Homebuyers” with “Why We Waived the Inspection & Saved Money,” which offers “unique expert or experienced takes that go beyond common knowledge.”

  • Primary sources over content farms

    Official

    Anthropic says early agents behind Claude’s Research mode “consistently chose SEO-optimized content farms over authoritative but less highly-ranked sources.” It fixed that with source-quality rules, and now grades the system on whether it used “primary sources over lower-quality secondary sources.”

  • Evidence raises visibility

    Observed

    In the 2023 research paper that named generative engine optimization, adding statistics, quotations and citations raised a source’s visibility in generated answers by up to about 40%.

  • Undercovered topics stay hot

    Official

    Perplexity says its index keeps documents from “undercovered topics” and authoritative domains fresh — a structural reward for covering what others haven’t.

The flip side is scaled content abuse: pages mass-produced to restate what already exists add nothing, and Google’s spam policies target exactly that purpose, whether the pages were written by people or by AI.

Rankbox framework

The Consensus Diff

A pre-publish check that turns information gain from an abstract idea into a count. Build the consensus the engines already have, subtract it from your draft, and see what’s left — if nothing is, the page is a rewrite.

  1. 01

    Build the consensus list

    Read the top five Google results and the AI answer to the page’s main question in ChatGPT or Perplexity. Write down every claim they share, one line each. This is what an engine can say without you.

  2. 02

    Diff your draft

    Mark every claim in your draft that’s already on the list. What stays unmarked is your gain. Zero unmarked claims means the page is a rewrite, however well it’s written.

  3. 03

    Fill the gap from what only you have

    Add at least one item to each major section from the five sources: first-party numbers, a test, first-hand experience, a named expert view, a worked example with real inputs.

  4. 04

    Lead with the gain

    Move the new fact into the first sentence of the section it belongs to, with a name, a number and a date, so it’s the part an engine lifts.

  5. 05

    Make it checkable

    Say how you know: the sample, the method, the date. An unsourced claim is easy to dismiss — for a reader, and for an engine choosing between sources.

How to use it: Aim for at least one unmarked, checkable claim in every major section. A section that can’t get one should be merged into another page or cut: it’s the part of the page with the least reason to be cited.

Free to use and adapt. If you cite it, link to rankbox.xyz/glossary/information-gain.

How do you add information gain to a page?#

Add information gain by finding what the top results and the AI answer already agree on, then contributing what they lack: numbers only your company has, a test you ran, a lesson from doing the work, an expert’s specific view, or a worked example with real inputs.

Source of gainExample for a SaaS teamWhy engines can use it
First-party numbersMedian setup time from your onboarding calls; the share of trial users who invite a teammateA specific, quotable figure no other page can supply
A test or teardownTiming five tools importing the same 500-task spreadsheetFirst-hand evidence — what Google calls non-commodity
ExperienceWhat broke when you migrated your own team, and the fixThe first E in E-E-A-T
A named expert viewYour lead engineer on when not to use a featureAn attributable claim with a source
A worked exampleThe actual cost for a 10-person team, line by lineSpecific enough to answer a narrow sub-query

Common mistakes with information gain#

The most common information gain mistakes are treating a longer article as a more original one, rewriting the top results in new words, and burying the one genuinely new finding deep in the page where neither readers nor AI engines reach it.

Myth

More words means more information.

Reality

Length isn’t gain. A 3,000-word page that restates the consensus adds less than one section with a new, sourced number.

Myth

Paraphrasing the top results makes content original.

Reality

Google’s guidance asks for “substantial additional value and originality” when a page draws on other sources. New wording isn’t new information.

Myth

Only big companies have original data.

Reality

Small teams have what big ones rarely publish: onboarding timings, support-ticket patterns, real quotes, mistakes made and fixed. One specific figure is enough to separate a page from the pack.

Myth

Save the original insight for the end.

Reality

Engines weight the top of a page. Growth Memo found 44.2% of ChatGPT citations came from the first 30% of the text, so lead with the new fact — the answer-first rule applied to your best material.

Sources

  1. 1.Contextual estimation of link information gain (US11354342B2)Google Patents · patents.google.com
  2. 2.Creating helpful, reliable, people-first contentGoogle Search Central · developers.google.com
  3. 3.Optimizing your website for generative AI featuresGoogle Search Central · developers.google.com
  4. 4.A guide to Google Search ranking systemsGoogle Search Central · developers.google.com
  5. 5.How we built our multi-agent research systemAnthropic Engineering · anthropic.com
  6. 6.Architecting and evaluating an AI-first search APIPerplexity Research · research.perplexity.ai
  7. 7.GEO: Generative Engine OptimizationAggarwal et al., KDD 2024 · arxiv.org
  8. 8.The science of how AI pays attentionGrowth Memo (Kevin Indig) · growth-memo.com

Know someone who’d find this useful? Send it their way.

Written by

Rankbox Team

The team behind Rankbox. We study how ChatGPT, Perplexity, Gemini, and Google AI Overviews choose their sources, and publish what we learn so you can put it to work.

See which AI answers cite you today

Enter your site to see how often ChatGPT, Perplexity, Gemini, and Google cite your brand, and exactly what to publish next.

No credit card required · Free 7-day trial