The "Shadow Training Data" Audit: How Common Crawl Decided Your Brand's Fate in 2024

A shadow training data audit maps AI model cutoffs to Common Crawl's 2024 crawls, checks what they held about your brand, and fixes the gaps.

Rankbox Team

September 29, 2026 · 22 min read

On this page10 sections

The short answer

A shadow training data audit checks what the open web said about your brand in the crawls that closed before each AI model's knowledge cutoff. Most AI labs no longer name their datasets, so Common Crawl, the free public web archive, is the closest record of that web you can still inspect. For a large group of models, the newest web they learned from dates from 2024. If your brand was thin or missing in those crawls, the model has little to remember, and only a live web search can introduce you.

The link between Common Crawl and training data used to be written down. OpenAI's GPT-3 paper gave a filtered Common Crawl 60% of its training mix, taken from 41 shards of monthly crawls covering 2016 to 2019. Meta's first LLaMA paper gave English Common Crawl 67%. Newer model cards talk about "publicly available" web data and name no crawl at all.

Memory still decides a lot of answers, because assistants often skip the search. In Profound's July 2026 test of about 400 prompts with web search switched on, Claude searched 36.6% of the time. Every other answer came from its training data.

This guide covers vendor disclosures, which crawls fall inside which model's window, how to query past crawls for your domain and for pages that named you, how to score the result, and what to do if you were invisible. For live lookups in today's crawl, Wikidata and Crunchbase, see our guide to entity authority in the AI era.

Key Takeaways

  • The clearest numbers on Common Crawl come from two papers: GPT-3 (60% of the training mix) and Meta's first LLaMA (67%, plus 15% from C4, a dataset also built from Common Crawl).
  • OpenAI, Anthropic, Google and Meta now describe their training data as some mix of public web, partner or licensed, user and synthetic data. Common Crawl is a public stand-in for the web their crawlers saw, not proof of what they used.
  • Models with knowledge cutoffs between April 2024 and January 2025 learned a web whose newest pages date from 2024 or early January 2025. That group includes GPT-5 (30 September 2024), Claude 3.7 Sonnet (October 2024) and Gemini 3 Pro (January 2025).
  • Common Crawl is a sample. Mistral AI's domain had 157 successful captures in the December 2024 crawl, and its Wikipedia article appeared in none of the eight crawls checked.
  • Pages that mention you can drop out of later crawls when their publisher blocks CCBot. Common Crawl's own copies show TechCrunch adding that block on 21 May 2024. Its 2023 story on Mistral's seed round appears in no later crawl checked, through September 2026.
  • The Shadow Audit Score rates each model window on presence, repetition, corroboration and fidelity, 0 to 2 each. The fictional Tallyfold scores 2, 5 and 6 out of 8 across three windows.
  • You can't edit training data after a model ships. Win the live search now, seed consistent entity facts, and make sure the next crawls capture your current story.

What AI Vendors Actually Disclose About Training Data

One claim you'll hear in AI SEO is that GPT-4o, Claude 3.5 and Gemini formed their view of your brand from years-old scrapes of Common Crawl, Wikipedia and Reddit. Part of that is documented and part of it is guesswork. Here is what each vendor has published about its training data, in its own papers and model cards.

VendorMost detailed disclosureWhat recent model cards sayNames Common Crawl?
OpenAIGPT-3 (2020): filtered Common Crawl, 410 billion tokens, 60% of the mix. WebText2 22%, two book sets 16%, Wikipedia 3%GPT-4o: "industry-standard machine learning datasets and web crawls," data up to October 2023. GPT-5: public internet data, partner data, and data from users, trainers and researchersGPT-3 only
AnthropicClaude 3 (2024): "publicly available information on the Internet as of August 2023," plus third-party, contractor and internal dataClaude Fable 5.1: public online information, public and private datasets, user data and synthetic dataNo. It describes its own crawler
GoogleGemini 3 Pro: "publicly-available web-documents," downloadable public datasets, "data obtained by crawlers," licensed and user dataProcessing includes "honoring robots.txt" and quality filteringNo
MetaLLaMA (2023): English Common Crawl 67% from five crawls, 2017 to 2020. C4 15%, Wikipedia 4.5%Llama 4: public, licensed and Meta product data, including public Instagram and Facebook posts. Muse Glimmer (August 2026): public, third-party and Meta product dataLLaMA 1 only

Sources: GPT-3, the GPT-4o system card, the GPT-5 system card, the Claude 3 model card, Anthropic's Transparency Hub, the Gemini 3 Pro model card, LLaMA, the Llama 4 model card and the Muse Glimmer model card.

What the Common Crawl story gets right

Web crawl was the largest slice in both disclosures that give numbers. Wikipedia sat in both mixes as its own source. Reddit shows up too: GPT-3's WebText2 is an expanded version of WebText, which OpenAI's GPT-2 paper built from "all outbound links from Reddit" that earned at least 3 karma.

Common Crawl also feeds open datasets that anyone can train on. Hugging Face's FineWeb holds more than 18.5 trillion tokens of English web text drawn from every Common Crawl dump since 2013. In January 2025 it added eight snapshots covering May to December 2024.

Common Crawl's Executive Director, Rich Skrenta, on TWiT's Intelligent Machines (August 2025). Watch on YouTube ↗

What it gets wrong

Current model cards don't name Common Crawl. Anthropic's model cards describe its own crawler, which follows robots.txt and "does not access password-protected or sign-in pages." Google lists "data obtained by crawlers" and licensed data. Meta's newest cards add data from its own apps. So the crawls you'll query below show what the public web looked like at a given moment. They don't show the training data any one lab collected.

"Years ago" is also out of date. As of September 2026, OpenAI lists a knowledge cutoff of 30 April 2026 for GPT-6 Astra, and Anthropic lists June 2026 for Claude Opus 5.5 on its models overview. The next section explains why 2024 still matters.

Captured is not the same as learned

Labs throw away much of what a crawl collects before it becomes training data. GPT-3's team trained a classifier with WebText, Wikipedia and books as examples of good text, then kept the Common Crawl documents that scored well. LLaMA's team trained a model to tell pages used as Wikipedia references from random pages, and "discarded pages not classified as references." Both also removed near-duplicates.

So a capture is the minimum, not the finish line. A thin homepage with 40 words of copy can sit in the crawl and still fall out of the training data.

Which Crawls Fell Inside Each Model's Training Window

A knowledge cutoff is the date a model's training data ends. Vendors publish it on model pages and model cards. Anthropic gives two dates: a training data cutoff and an earlier "reliable knowledge cutoff," the point up to which knowledge is "most extensive and reliable." Claude Haiku 4.5, for example, lists February 2025 and July 2025.

Treat any published date as approximate. The Dated Data study found that "effective cutoffs often differ from reported cutoffs," partly because new Common Crawl dumps contain "non-trivial amounts of old data." The last months before a cutoff are thin in the training data, likely because the web hasn't finished writing about them yet.

With that caveat, you can line each model up against Common Crawl's calendar. We call the result the Crawl-Window Map. Crawl dates come from Common Crawl's index list, and cutoffs from each vendor's own page.

ModelVendor-stated cutoffLast crawl that ended before it2024 crawls inside the window
GPT-4o1 October 2023CC-MAIN-2023-23 (2023-40 straddles it)0
Claude 3.5 SonnetApril 2024CC-MAIN-2024-182
GPT-5 mini31 May 2024CC-MAIN-2024-223
GPT-4.1 and o31 June 2024CC-MAIN-2024-223
Claude 3.5 HaikuJuly 2024CC-MAIN-2024-305
Gemini 2.0 Flash and Llama 4August 2024CC-MAIN-2024-336
GPT-5 and GPT-5.130 September 2024CC-MAIN-2024-387
Claude 3.7 SonnetOctober 2024CC-MAIN-2024-428
Gemini 2.5 Pro and Gemini 3 ProJanuary 2025CC-MAIN-2025-0510
Newest flagships, September 2026January to June 2026A late 2025 or 2026 crawl10, plus later crawls

The cutoffs come from OpenAI's model pages (for example GPT-5), Anthropic's Claude 3.5 addendum and Transparency Hub, Google's model pages for Gemini 2.0 Flash and Gemini 2.5 Pro, and Meta's model cards.

Why the title says 2024

Common Crawl ran ten crawls in 2024, from CC-MAIN-2024-10 (20 February to 5 March) to CC-MAIN-2024-51 (1 to 15 December). Every model from Claude 3.5 Sonnet down to Gemini 3 Pro learned from a web that stopped between April 2024 and January 2025. Some of that memory is still in service, and some of it lives on inside newer models:

  • OpenAI still lists GPT-5 as a "previous model" in its API, with its 30 September 2024 cutoff.
  • Google's model cards for Gemini 3.7 Flash and Gemini 3.8 Flash give a March 2026 cutoff, then warn that in some domains "the model's knowledge is limited to January 2025 (in line with the Gemini 3 Model Family)."
  • Anthropic's current lineup includes Claude Haiku 4.5, with a reliable knowledge cutoff of February 2025.

So for many models people used through 2025, and part of what they use now, the newest web in their training data is from 2024. A brand that launched, pivoted or got its first press in 2024 was either caught by those crawls or missed.

How to Query Historical Common Crawl Indexes for Your Domain

Common Crawl keeps a separate URL index for every crawl at index.commoncrawl.org. Queries are free and need no key, and they show which pages could have become training data in each period. The examples use Mistral AI, a real company whose records are public, because its timing makes a clean case: Mistral says on its About page that it was born in April 2023, hired its first employee on 5 June and closed its seed round on 13 June. Everything else in this guide uses Tallyfold, a made-up company.

Every request below was run on 29 September 2026.

Step 1: List the crawls and their dates

bash
curl -s "https://index.commoncrawl.org/collinfo.json"

On 29 September 2026 the list held 128 crawls, newest first. Each entry, trimmed here, gives an ID, a query endpoint and the dates the crawl ran:

json
{
"id": "CC-MAIN-2024-10",
"name": "February/March 2024 Index",
"cdx-api": "https://index.commoncrawl.org/CC-MAIN-2024-10-index",
"from": "2024-02-20T21:10:55",
"to": "2024-03-05T15:40:45"
}

The ID reads as year and week. Match the to date against each model's cutoff to build your own window list.

Step 2: Ask the crawls in each window for your domain

Start with the crawl just before your launch and move forward. Send a descriptive user agent with a contact address:

bash
curl -s -A "AcmeShadowAudit/1.0 (you@example.com)" \
"https://index.commoncrawl.org/CC-MAIN-2023-23-index?url=mistral.ai&matchType=domain&output=json"

The May/June 2023 crawl, which ended on 11 June 2023, answered with HTTP 404 and {"message": "No Captures found for: mistral.ai"}. The next crawl, CC-MAIN-2023-40, ran from 21 September to 5 October 2023. It returned five captures. The one that matters, trimmed:

json
{
"urlkey": "ai,mistral)/",
"timestamp": "20231001221357",
"url": "https://mistral.ai/",
"status": "200",
"length": "5240",
"offset": "420680499",
"filename": "crawl-data/CC-MAIN-2023-40/segments/1695233510941.58/warc/CC-MAIN-20231001205332-20231001235332-00601.warc.gz"
}

The urlkey is your domain written backwards, timestamp is when CCBot fetched the page, and status 200 means it got the page itself. The last three fields locate the stored copy. That capture landed on 1 October 2023, four days after Mistral released its first model and on GPT-4o's cutoff date.

Step 3: Count captures crawl by crawl

One number per crawl tells you whether your share of the open training data pool was growing. This loop counts successful captures across 2024 and pauses between calls:

bash
for c in CC-MAIN-2024-10 CC-MAIN-2024-18 CC-MAIN-2024-22 CC-MAIN-2024-26 \
CC-MAIN-2024-30 CC-MAIN-2024-33 CC-MAIN-2024-38 CC-MAIN-2024-42 \
CC-MAIN-2024-46 CC-MAIN-2024-51; do
n=$(curl -s -A "AcmeShadowAudit/1.0 (you@example.com)" \
"https://index.commoncrawl.org/$c-index?url=yourdomain.com&matchType=domain&output=json&fl=url,status" \
| grep -c '"status": "200"')
echo "$c $n"
sleep 5
done

For a large site, add &showNumPages=true first and fetch each page with &page=N. Here is Mistral AI's domain, subdomains included:

CrawlDatesCapturesStatus 200Hostnames
CC-MAIN-2023-2327 May–11 Jun 2023000
CC-MAIN-2023-4021 Sep–5 Oct 2023521
CC-MAIN-2023-5028 Nov–12 Dec 202343293
CC-MAIN-2024-1020 Feb–5 Mar 202479484
CC-MAIN-2024-1812–25 Apr 2024112545
CC-MAIN-2024-2217–31 May 2024125509
CC-MAIN-2024-2612–25 Jun 202493396
CC-MAIN-2024-3012–25 Jul 2024179868
CC-MAIN-2024-332–16 Aug 2024187968
CC-MAIN-2024-387–21 Sep 2024181869
CC-MAIN-2024-423–16 Oct 202424812710
CC-MAIN-2024-461–15 Nov 20242511638
CC-MAIN-2024-511–15 Dec 202426115710

Three things stand out. First, the numbers are small for a company whose seed round made news around the world. Common Crawl's FAQ says its dataset "is a sample of the web, and we do not generally archive any entire website but a randomly selected subset of it."

Second, timing is luck. Mistral announced Mistral Large on 26 February 2024. CC-MAIN-2024-10 fetched the homepage the next day but missed the announcement, which first appears on 24 April 2024 in CC-MAIN-2024-18. Third, even a homepage can go missing: CC-MAIN-2024-42 holds 127 good captures from the domain, and none is the homepage.

Step 4: Read the copy the crawl stored

The index points at a byte range inside a large WARC file, the archive format Common Crawl stores pages in. Ask for just that range, using offset as the start and offset + length - 1 as the end:

bash
curl -s -r 420680499-420685738 \
"https://data.commoncrawl.org/crawl-data/CC-MAIN-2023-40/segments/1695233510941.58/warc/CC-MAIN-20231001205332-20231001235332-00601.warc.gz" \
| gunzip | grep -o "<title>[^<]*</title>"

It prints <title>Mistral AI | Open source models</title>. Remove the grep to see the whole page. The copy, as stored on 1 October 2023, says "Our teaser model is out! The best 7B, Apache 2.0." and "Mistral 7B is better than Llama 2 13B on all benchmarks." The meta description reads "Frontier AI in your hands."

That page is Mistral's shadow training data. Any system that learned from that crawl met a small open-model lab with one 7B model. Its later models, funding rounds and customers were not on the page yet. It's the easiest step to skip, and the one that tells you what a model could have learned, not just whether you existed.

Be polite to the index server

Common Crawl's FAQ says the index API is "heavily rate limited," asks you to sleep between calls, and warns that a blocked IP should wait 24 hours. For bulk work, it points to its URL Index in Amazon Athena or Apache Spark. On 29 September 2026 the server stopped answering partway through this audit, so the crawl-by-crawl counts above and the TechCrunch lookups below were read from the same index files, which Common Crawl also publishes on data.commoncrawl.org.

How to Find the Pages That Mentioned You

Your own site is half of your shadow training data. The other half is what others wrote about you. Common Crawl's index lists URLs, not words, so you can't search it for your brand name. There are three routes, in order of effort.

  1. Check the URLs you already know. List your press coverage, review profiles, directory listings and partner pages, then query each exact URL in each crawl in the window.
  2. Scan a URL prefix and filter it. News sites put the story's name in the URL. Ask for a month of a publisher's URLs and filter for your brand.
  3. Read the text. For each candidate URL, fetch the stored copy as in Step 4 and check how it describes you.

Here's route 2 for the TechCrunch story on Mistral's seed round, which ran on 13 June 2023:

bash
curl -s -A "AcmeShadowAudit/1.0 (you@example.com)" \
"https://index.commoncrawl.org/CC-MAIN-2024-10-index?url=techcrunch.com/2023/06/&matchType=prefix&output=json&fl=timestamp,url,status" \
| grep -i mistral

It returns one line: the article, fetched on 22 February 2024 with status 200. Repeat the check across crawls to trace the story's path. It first shows up in CC-MAIN-2023-40, carrying newsletter tracking parameters. CC-MAIN-2023-50 holds it too, alongside a second TechCrunch piece on Mistral from 16 June. CC-MAIN-2024-18 missed it, and CC-MAIN-2024-22 caught it three more times, on 18 and 21 May 2024. Then it's gone.

Why mentions vanish from later crawls

The reason sits in TechCrunch's robots.txt, which Common Crawl also stores. The May 2024 crawl fetched that file 100 times. The copies up to 15:33 UTC on 21 May have no rule for CCBot. From 21:19 UTC that day, every copy contains User-agent: CCBot and Disallow: /, next to the same rule for GPTBot. None of the later crawls checked, from June 2024 to September 2026, holds a single TechCrunch URL from June 2023.

TechCrunch isn't unusual. The Consent in Crisis audit of 14,000 web domains found that in one year, 2023 to 2024, new restrictions left 5% or more of all tokens in C4, and 28% or more of its most actively maintained critical sources, "fully restricted from use." So your press footprint in open training data can shrink even while every article stays online. Publishers choose whom to allow, and the choice differs by crawler. Our AI crawler directory lists the tokens they use.

Why Wikipedia looks missing

A Wikipedia page missing from Common Crawl isn't missing from the training data. The December 2024 crawl held 61,328 captures from en.wikipedia.org, against more than 7.2 million English articles today, and Mistral AI's article appeared in none of the eight 2023 and 2024 crawls checked. Labs add Wikipedia separately: 3% of GPT-3's mix, and 4.5% of LLaMA's from Wikipedia dumps. Check Wikipedia and Wikidata directly, as the entity authority guide shows.

The Shadow Audit Score: Reading Your Results

Raw counts don't say much on their own. The Shadow Audit Score turns them into a verdict on your shadow training data for each model window you care about. Pick three models your buyers actually use, look up their cutoffs on the vendors' pages, and score each window on four checks.

Check0 points1 point2 points
PresenceNo status 200 capture of your homepage in the windowHomepage onlyHomepage plus at least two of About, product and pricing pages
RepetitionNo crawl in the window captured you1 or 2 crawls did3 or more crawls did
CorroborationNo third-party URL naming you in window crawls1 to 4 distinct URLs5 or more distinct URLs
FidelityNo captured homepage copy matches today's one-line descriptionFewer than half of the captured homepages matchHalf or more match

Add the four scores for a total out of 8, then read it as a band:

  • 0 to 2, Invisible. The model's training data held almost nothing about you. Expect "I don't have information about that company," a guess, or a namesake. Call it entity invisibility. It's normal for a brand born after the cutoff.
  • 3 to 5, Faint. The model may know your name but not much else, or know an old version of you.
  • 6 to 8, Present. The web in that window described you clearly and more than once. Check fidelity before you relax.

Fidelity carries extra weight because repetition shapes memory. Kandpal and colleagues showed that a model's ability to answer a factual question tracks how many training documents about it the model saw. If most stored copies describe your old product, the old product is the likelier answer. Presence, repetition and corroboration explain whether a model could know you. Fidelity explains what it would say.

Worked Example: Tallyfold's Shadow Audit

Tallyfold is a made-up company, and so is every number in this section. It sells invoicing and payments software to agencies, at the reserved domain tallyfold.example. The story is invented for illustration. Tallyfold launched its site in mid-March 2024 as "payments for freelancers," then repositioned as "invoicing and payments software for agencies" in October 2024.

Here is what its shadow training data audit might return:

CrawlStatus 200 capturesHomepage captured?Homepage copyThird-party URLs naming Tallyfold
CC-MAIN-2024-100NoSite not live0
CC-MAIN-2024-181YesOld: freelancers0
CC-MAIN-2024-220NoNone0
CC-MAIN-2024-263Yes, plus pricing and featuresOld0
CC-MAIN-2024-302YesOld0
CC-MAIN-2024-334YesOld1 partner directory page
CC-MAIN-2024-385YesOld0
CC-MAIN-2024-420NoNone0
CC-MAIN-2024-466YesNew: agencies1 podcast episode page
CC-MAIN-2024-517YesNew0
CC-MAIN-2025-058YesNew1 review listing

Now score three windows, using cutoffs from the Crawl-Window Map:

CheckClaude 3.5 Sonnet (April 2024)GPT-5 (30 September 2024)Gemini 3 Pro (January 2025)
Crawls in the window2024-10, 2024-182024-10 to 2024-382024-10 to 2025-05
Presence1 (homepage only)2 (pricing and features from 2024-26)2
Repetition1 (1 crawl)2 (5 crawls)2 (8 crawls)
Corroboration01 (1 URL)1 (3 URLs)
Fidelity0 (0 of 1 homepages)0 (0 of 5)1 (3 of 8)
Total out of 82, Invisible5, Faint6, Present

Check the arithmetic. The GPT-5 window holds seven crawls, and five of them captured something: 2024-18, -26, -30, -33 and -38. That's 1 + 3 + 2 + 4 + 5 = 15 captures. The Gemini 3 Pro window adds four more crawls, three with captures: 15 + 0 + 6 + 7 + 8 = 36 captures from 8 crawls. Three of those eight homepages carry the new copy, which is 37.5%, under half, so fidelity scores 1.

Here's what each result means for Tallyfold's buyers.

  • Claude 3.5 Sonnet window: invisible. From memory, such a model knows little beyond a name. A useful answer needs a search.
  • GPT-5 window: faint, and wrong. Everything it could have seen says "freelancers." A confident old answer is harder to fix than no answer.
  • Gemini 3 Pro window: present but mixed. The new story is there, outnumbered five crawls to three, so answers may blend both.

Tallyfold can't change the training data behind any of these windows. What it can change is the next one, and what search-grounded answers read today. For the retrieval side of a pivot like this, see our guide to semantic drift after a product pivot.

What to Do If You Were Invisible in 2024

No page you publish reaches into a trained model. Only the vendor can change its weights, by training a new model on new training data. That leaves three tracks, and each runs on its own clock.

Track 1: Win the live search (weeks)

A search is where an assistant can meet you for the first time. Anthropic's web search docs say Claude searches for information "beyond its knowledge cutoff," including "information about specific organizations, people, or products that might have changed." Profound saw words like "best," "near me" and the current year pull Claude into a search, while basic "what is" prompts were more likely to stay inside the model.

  1. Make sure search crawlers can reach you. Allow OAI-SearchBot, Claude-SearchBot, PerplexityBot and the Google and Bing crawlers. Test your file with the free robots.txt tester.
  2. Publish a page that answers "What is [brand]?" in its first sentence, with your category, audience, founding date and pricing.
  3. Target the prompts that trigger search: comparisons, "best X for Y," current pricing and alternatives. Our breakdown of how AI models rank brands in search results covers what happens after the search step.

Track 2: Seed structured entity facts (weeks to months)

Structured entity seeding means stating the same facts, the same way, everywhere machines look. It helps retrieval now and gives the next crawl one story to capture.

  1. Write one canonical definition sentence and use it on your homepage, About page and every profile you control.
  2. Add Organization markup with a stable @id and sameAs links to your profiles. The free schema generator gives you a starting block, and our guide to building a knowledge graph for AI covers the patterns.
  3. Earn independent pages that name you next to your category. Check each publisher's robots.txt: a site that blocks CCBot stays out of Common Crawl and the open datasets built on it, though its pages still help live search.
  4. Treat Wikidata as a result of coverage, not a tactic. What entity authority in SEO means explains why independent sources matter more than self-made records.

Track 3: Get captured by the next crawls (months)

The training data for the next models is being crawled now. If you want to be learned, make it easy.

  1. Decide on purpose whether to allow training crawlers such as CCBot, GPTBot and ClaudeBot, and the Google-Extended token. Blocking them keeps your own pages out of those crawls and the LLM training data built from them. Third-party pages about you can still get in. The AI robots.txt generator writes either choice.
  2. Announce your sitemap in robots.txt. Common Crawl's FAQ says its crawler uses "any Sitemap announced in the robots.txt file."
  3. Link your About, product and pricing pages from the homepage, and keep their URLs stable across redesigns.
  4. Retire old positioning everywhere you control, so the next crawl stores one story, not two.
  5. Re-run the audit when a new crawl appears in collinfo.json, and score each new model's window once its cutoff is published.
Your bandThis monthFor the next model
Invisible (0–2)Track 1 first: answer "What is [brand]?" on a crawlable pageTracks 2 and 3 in full
Faint (3–5)Track 1, plus fix the facts retrieval findsRaise repetition and corroboration
Present, low fidelityRetire the old story on your site and profilesMake the new copy the majority in the next crawls
Present, high fidelityKeep facts consistentKeep crawlers allowed and re-audit twice a year

If AI answers already state wrong facts about you, our guide to fixing incorrect brand facts in AI answers has the full recovery plan.

Where Rankbox Fits in a Shadow Training Data Audit

Rankbox doesn't query Common Crawl, audit past crawls or put anything into a model's training data. Nobody can do that last part. It also doesn't track AI citations today. What it does is the writing half of Tracks 1 and 2. The Citation-Ready Writer researches the live web and writes 2,000 to 3,500-word source-backed articles. Brand Voice applies the tone, audience, style rules and product details you give it, so every article uses your canonical one-liner and mentions your product where it fits.

Articles reach your site through the Rankbox API, which a developer wires in. Each one is another crawlable page that states your current story. The Business plan is $49.50 a month with a 7-day trial. See pricing.

Frequently Asked Questions

What is shadow training data?

Shadow training data is the version of your brand that sat in web crawls before a model's knowledge cutoff. Vendors don't publish their exact datasets, so you can't see it directly. Common Crawl's historical indexes are the closest public record, because they show which of your pages, and which pages about you, were collected in each period and what they said.

Did ChatGPT, Claude and Gemini train on Common Crawl?

No vendor confirms it for current models. OpenAI's 2020 GPT-3 paper, a predecessor of the models behind ChatGPT, gave filtered Common Crawl 60% of its training data mix. Later OpenAI, Anthropic and Google model cards describe public web data, their own crawlers and licensed data without naming Common Crawl. Treat it as a stand-in for the public web those crawlers saw, not as a confirmed ingredient.

How do I check if my website was in Common Crawl in 2024?

Query each 2024 crawl's index at index.commoncrawl.org with your domain. Use crawl IDs CC-MAIN-2024-10 through CC-MAIN-2024-51, add matchType=domain&output=json, and count lines with status 200. A 404 with "No Captures found" means that crawl didn't collect your site. Sleep a few seconds between requests.

What happens if my brand was founded after a model's knowledge cutoff?

The model has no memory of you. From training data alone it can only guess, decline, or confuse you with a namesake. It can still describe you correctly when it runs a web search and finds a clear page about you. That's why crawlable, search-ready pages matter most for young brands.

Can I get my brand added to an AI model's training data?

Not directly. There's no submission form, and trained weights can't be edited from outside. You can make your site crawlable by training bots, keep your facts consistent, earn independent coverage, and let the next crawls capture that. Whether a lab uses those pages is its decision.

Should I block CCBot?

Block CCBot only if keeping your content out of open datasets matters more than being learned. Common Crawl's archive feeds public datasets such as FineWeb. Blocking it doesn't affect live AI search, which uses other crawlers. Most brands that want to be recommended allow it and decide on other training crawlers one by one.

References

  1. 1.Language Models are Few-Shot Learners (Brown et al., 2020), arXivarxiv.org ↗
  2. 2.Language Models are Unsupervised Multitask Learners (Radford et al., 2019), OpenAIcdn.openai.com ↗
  3. 3.GPT-4o System Card, OpenAI via arXivarxiv.org ↗
  4. 4.GPT-5 System Card, OpenAIcdn.openai.com ↗
  5. 5.GPT-5 model page, OpenAI API docsdevelopers.openai.com ↗
  6. 6.The Claude 3 Model Family: Opus, Sonnet, Haiku (model card), Anthropicassets.anthropic.com ↗
  7. 7.Model Card Addendum: Claude 3.5 Haiku and Upgraded Claude 3.5 Sonnet, Anthropicassets.anthropic.com ↗
  8. 8.Transparency Hub, Anthropicanthropic.com ↗
  9. 9.Models overview, Claude Platform docsplatform.claude.com ↗
  10. 10.Gemini 3 Pro Model Card, Google DeepMindstorage.googleapis.com ↗
  11. 11.Gemini 3.8 Flash model card, Google DeepMinddeepmind.google ↗
  12. 12.Gemini 2.0 Flash, Gemini API docsai.google.dev ↗
  13. 13.LLaMA: Open and Efficient Foundation Language Models (Touvron et al., 2023), arXivarxiv.org ↗
  14. 14.Llama 4 model card, Meta on GitHubgithub.com ↗
  15. 15.Muse Glimmer model card, Meta on Hugging Facehuggingface.co ↗
  16. 16.Common Crawl Index Server collection list, Common Crawlindex.commoncrawl.org ↗
  17. 17.FAQ, Common Crawlcommoncrawl.org ↗
  18. 18.Dated Data: Tracing Knowledge Cutoffs in Large Language Models (Cheng et al., 2024), arXivarxiv.org ↗
  19. 19.Consent in Crisis: The Rapid Decline of the AI Data Commons (Longpre et al., 2024), arXivarxiv.org ↗
  20. 20.FineWeb dataset card, Hugging Facehuggingface.co ↗

Written by

Rankbox Team

The team behind Rankbox. We study how ChatGPT, Perplexity, Gemini, and Google AI Overviews choose their sources, and publish what we learn so you can put it to work.

Who we are and how we work

See where AI cites you today

Enter your site to see how often ChatGPT, Perplexity, Gemini, and Google cite your brand, and exactly what to publish next.

No credit card required · Free 7-day trial