AI hallucination

nounalso called LLM hallucination or confabulation

Definition

An AI hallucination is a confident but false statement produced by a large language model — an invented statistic, feature, price or citation — which happens because the model generates plausible text rather than looking facts up.

Updated 5 min read5 cited sources

On this page7 sections

Why it matters for founders and small teams

When an AI assistant tells a buyer your product lacks a feature it has, or quotes last year’s price, you lose the deal without ever knowing it happened. Small brands are the most exposed, because the facts a model is most likely to invent are the ones that appear rarely on the web — which describes most facts about a young company.

Why do AI models hallucinate?#

AI models hallucinate because a large language model generates the most plausible next words rather than looking facts up, and training and evaluation reward a confident guess over admitting uncertainty — so rare facts get filled in with fluent fiction.

OpenAI researchers made the case in a September 2025 paper, Why Language Models Hallucinate: “Like students facing hard exam questions, large language models sometimes guess when uncertain, producing plausible yet incorrect statements instead of admitting uncertainty.” Benchmarks that score only right or wrong answers teach models that a guess beats “I don’t know.”

The paper’s most useful point for brands is about rare facts. Facts with no learnable pattern — a birthday, a founding date, a price — can only be memorized, and the authors estimate that “if 20% of birthday facts appear exactly once in the pretraining data, then one expects base models to hallucinate on at least 20% of birthday facts.” A fact about your company that appears on one page of the web is a fact a model is likely to get wrong.

Can AI hallucinate about my company?#

AI can hallucinate about any company, and the most common brand hallucinations are outdated prices, invented or missing features, wrong integrations, confusion with similarly named companies and fabricated citations — usually when the model answers from memory or from a thin, ambiguous source.

TypeExampleUsual cause
Stale factQuotes a price you changed in MarchAnswering from training data older than the change — see knowledge cutoff
Invented detailSays you offer a free plan you’ve never hadA rare fact filled in with a plausible guess
Entity confusionMixes you up with a similarly named companyWeak or inconsistent entity signals — see entity SEO
MisattributionCites a reseller’s outdated page for your specsRetrieval found a third-party copy before your page
Fabricated sourceLinks to a URL that doesn’t existThe model generating a citation instead of retrieving one

Fabricated sources are not rare: the Tow Center study found that more than half of Gemini’s and Grok 3’s responses cited fabricated or broken URLs. And because Gemini names brands in text far more often than it links them — only 21.4% of its brand appearances carried a link in a June 2026 Semrush study — an unlinked claim about you may have no source you can trace.

Worked example

The Brand Accuracy Audit

A way to turn “ChatGPT says weird things about us” into a number you can track and a short list of fixes. The inputs are illustrative, for a fictional project tool called Plannora.

  1. 1

    Prompts in the panel

    25 buyer questions about Plannora: pricing, features, integrations and comparisons.

    25 prompts

  2. 2

    Engines and runs

    Each prompt run once in ChatGPT, Perplexity, Claude and Gemini.

    100 answers

  3. 3

    Answers that mention Plannora

    Only answers that talk about the brand can be wrong about it.

    72 answers

  4. 4

    Answers with at least one false claim

    Scored against the fact sheet: 9 quote the old $12 price, 4 invent a free plan, 3 miss the Slack integration launched in June.

    16 answers

  5. =

    Brand error rate

    16 ÷ 72 — the share of brand mentions that carry a false claim.

    22%

The result: The audit points at fixes, not just a score. In this example, 9 of the 16 errors trace back to one reseller’s pricing page cited by two engines, so one email and a clearer official pricing page address more than half the problem. Re-run the same panel once the corrected pages have been recrawled and compare the rate.

Free to use and adapt. If you cite it, link to rankbox.xyz/glossary/ai-hallucination.

How do you measure AI hallucinations about your brand?#

Measure brand hallucinations with a fixed panel of prompts about your company, run in each engine on a schedule, with every factual claim checked against a single source of truth — then track the share of brand mentions that contain at least one false claim.

  1. Write your fact sheet first: pricing, plans, core features, integrations, founding date, locations and who the product is for. This is the answer key.
  2. Build 20–30 prompts buyers really ask: “How much does Plannora cost?”, “Does Plannora integrate with Slack?”, “Plannora vs Loopcraft for a small team.”
  3. Run them in each engine, with search on and off where the product lets you. The difference shows whether an error comes from training or from a retrieved page.
  4. Score each answer against the fact sheet, and log the cited source behind every wrong claim.
  5. Repeat monthly and after every price or product change. Answers vary from run to run, so trends across several runs matter more than any single error. This is prompt tracking with an accuracy column.

How do you fix AI hallucinations about your brand?#

Fix brand hallucinations by publishing the correct facts where retrieval will find them — clearly titled, dated pages on your own domain, stated in plain sentences — and by correcting the third-party pages engines cite, since nobody can edit a model’s memory directly.

  • Publish an official page for every checkable fact: pricing, plans, integrations, security, comparisons. Since August 2026 most ChatGPT fan-out searches use site: to check specific domains, often the vendor’s own, so an official page is frequently the first place it looks.
  • State facts in citable sentences. “Plannora’s Team plan costs $8 per user per month, billed annually” beats a pricing grid rendered by JavaScript that no AI fetcher except Google’s can read.
  • Date and update honestly. Show a visible updated date and change it only with real changes. See content freshness.
  • Fix the sources engines cite. Ask review sites, directories and resellers to update stale listings, and require canonicals on syndicated copies of your content.
  • Make the entity unambiguous. Consistent naming, a clear About page, and Organization schema with sameAs links help engines tell you apart from similarly named companies.
  • Flag wrong answers with each engine’s feedback buttons. Vendors don’t say how that feedback is used, so treat it as a supplement, not a fix.

Sources

  1. 1.Why Language Models HallucinateKalai et al., OpenAI, 2025 · arxiv.org
  2. 2.We compared eight AI search engines. They're all bad at citing newsColumbia Journalism Review, Tow Center · cjr.org
  3. 3.The ghost citations studySemrush · semrush.com
  4. 4.ChatGPT tripled its fan-out queriesNectiv · nectivdigital.com
  5. 5.Web fetch toolClaude Developer Platform · platform.claude.com

Know someone who’d find this useful? Send it their way.

Written by

Rankbox Team

The team behind Rankbox. We study how ChatGPT, Perplexity, Gemini, and Google AI Overviews choose their sources, and publish what we learn so you can put it to work.

See which AI answers cite you today

Enter your site to see how often ChatGPT, Perplexity, Gemini, and Google cite your brand, and exactly what to publish next.

No credit card required · Free 7-day trial