Friday, Jul 24, 2026 The claims desk. Receipts included. POWERED BY LENZ
IsThis

TECH

The Claim

AI language models generate hallucinated or factually incorrect outputs in more than 20% of cases.

The Short Version

Hallucination rates above 20% are documented in specific high-stakes domains like medical literature review and clinical decision support, but the claim's unqualified framing suggests this is typical across all AI language model use — which the evidence does not support. Broad benchmarks show top current models averaging under 10%, and sometimes below 1%. The rate varies dramatically by model, task, domain, and how "hallucination" is measured, making a single blanket figure misleading.

Caveats

  • The >20% figures cited in supporting studies come primarily from specialized medical/clinical tasks and older models (e.g., GPT-3.5, Bard), not general-use scenarios.
  • Hallucination rates are highly dependent on task type, model generation, prompting method, and evaluation metric — 'cases' is undefined in the claim and could refer to prompts, answers, tokens, or citations.
  • Broad benchmark data (Vectara Hallucination Leaderboard, Frontiers survey) shows current top models averaging well under 10% on factual consistency tasks, directly contradicting the >20% threshold as a general rule.

The Receipts

  1. vectara/hallucination-leaderboard - GitHub

    Vectara Hallucination Leaderboard

  2. Hallucination Rates and Reference Accuracy of ChatGPT and Bard ...

    PubMed Central

  3. Multi-model assurance analysis showing large language models are highly vulnerable to adversarial hallucination attacks during clinical decision support - PMC

    PMC

  4. Survey and analysis of hallucinations in large language models

    Frontiers in Artificial Intelligence

  5. Evaluating the Accuracy of Responses by Large Language Models for Information on Disease Epidemiology - PMC

    PMC

  6. Survey of Hallucination in Natural Language Generation (Updated 2025)

    arXiv

  7. AI Hallucination: Compare top LLMs like GPT-5.2 - AIMultiple

    AIMultiple

  8. Are AI Hallucinations Getting Better or Worse? We Analyzed the Data | ScottGraffius.com

    ScottGraffius.com

  9. Reliability for unreliable LLMs - The Stack Overflow Blog

    The Stack Overflow Blog

  10. Study: Heavy AI Users See 3x More Hallucinations - Rev

    Rev

+ 1 more sources — see the full list on Lenz

Filed Under

AI language models

More Fact Checks