Sunday, Sep 6, 2026 The claims desk. Receipts included. POWERED BY LENZ
IsThis

AI error rates vary widely across models, tasks, and benchmarks.

The Claim

Artificial intelligence systems produce incorrect answers in 70% of evaluated cases.

The Short Version

Available evidence does not establish a general 70% error rate for artificial intelligence systems. Reported rates vary dramatically by model, task, benchmark, and definition of error. A result near 70% appears in a narrow medical evaluation, but presenting it as broadly representative of evaluated AI answers is unsupported.

Caveats

  • The 70% figure appears to come from a narrow evaluation and cannot be generalized to all AI systems.
  • Error and hallucination rates depend heavily on the task, model, benchmark, and scoring method.
  • Some apparent corroboration comes from secondary headlines lacking enough methodological context to verify equivalence.

The Receipts

  1. Evaluation of ChatGPT as a diagnostic tool for medical learners and clinicians

    journals.plos.org

  2. Evaluating the AI Potential as a Safety Net for Diagnosis: A Novel Benchmark of Large Language Models in Correcting Diagnostic Errors

    medrxiv.org

  3. arxiv.org

    arxiv.org

  4. Reliability without Validity: A Systematic, Large-Scale Evaluationof LLM-as-a-Judge Models Across Agreement, Consistency, and Bias

    arxiv.org

  5. AI chatbots provide poor answers to medical questions half the time ...

    cidrap.umn.edu

  6. AI Search Has a Citation Problem - Columbia Journalism Review

    cjr.org

  7. Google finds AI chatbots are only 69% accurate… at best - Digital Trends

    digitaltrends.com

  8. Evaluation and mitigation of the limitations of large language models in clinical decision-making | Nature Medicine

    preview-www.nature.com

  9. Dissecting clinical reasoning failures in frontier artificial intelligence using 10,000 synthetic cases

    medrxiv.org

  10. Responsible AI | The 2026 AI Index Report - Stanford HAI

    hai.stanford.edu

+ 19 more sources — see the full list on Lenz

Filed Under

Artificial Intelligence Systems

More Fact Checks