Friday, Jul 24, 2026 The claims desk. Receipts included. POWERED BY LENZ
IsThis

TECH

The Claim

Substantive disagreements between AI models on fact-checking outcomes are common.

The Short Version

Evidence from multiple studies shows that AI fact-checking models often reach materially different verdicts on the same claim, with reported substantive conflicts commonly in the roughly 15% to 30% range on challenging datasets. That is frequent enough to count as common in real-world use. Rates do vary by claim difficulty, ambiguity, prompting, and evidence quality.

Caveats

  • 'Common' should not be read as 'most claims' in every domain; disagreement is concentrated on harder, ambiguous, or politically contentious claims.
  • Some studies showing different overall label distributions do not, by themselves, prove claim-by-claim disagreement; the strongest evidence comes from direct pairwise conflict analyses.
  • Disagreement rates are sensitive to experimental setup, including prompt design, provided evidence, language, and benchmark composition.

The Receipts

  1. Efficient Annotator Reliability Assessment and Sample Weighting for ...

    arXiv

  2. Assessing Inter-Annotator Agreement for Medical Image Segmentation

    PubMed Central

  3. Towards Automated Fact-Checking of Real-World Claims

    ROMCIR / University of Milano-Bicocca

  4. The perils and promises of fact-checking with large language models

    NPJ Digital Medicine (via PubMed Central)

  5. A Multilingual, Comparative Analysis of LLM-Based Fact-Checking From Check-Worthiness to Verdict

    arXiv

  6. Cross-checking journalistic fact-checkers: The role of sampling and rating scales for estimating interrater reliability

    PubMed Central (Journalism & Mass Communication Quarterly)

  7. A Survey on Automated Fact-Checking

    Transactions of the Association for Computational Linguistics

  8. Unveiling Pitfalls and Potentials in Fact Verifiers

    OpenReview

  9. “Fact-checking” fact checkers: A data-driven approach

    Harvard Kennedy School Misinformation Review

  10. Fine-Grained Evaluation Benchmark for Automatic Fact-checkers

    arXiv

+ 18 more sources — see the full list on Lenz

Filed Under

Artificial Intelligence Models

More Fact Checks