Friday, Sep 11, 2026 The claims desk. Receipts included. POWERED BY LENZ
IsThis

As of Q1 2026, benchmarks don’t show frontier coding models beating experts.

The Claim

As of Q1 2026, frontier AI coding models exceed expert human performance on real-world software engineering tasks, as demonstrated by SWE-bench Verified and HumanEval+ results.

The Short Version

Available evidence does not show that frontier AI coding models outperform expert humans on real-world software engineering as of Q1 2026. Very high scores on SWE-bench Verified and HumanEval+ are not direct expert-versus-model comparisons, and HumanEval+ is a weak proxy for real software engineering. Independent analyses also report contamination, benchmark artifacts, and many supposedly successful patches that human maintainers would reject.

Caveats

  • A high benchmark score is not the same as beating expert humans unless both are measured on the same tasks under the same conditions.
  • HumanEval+ mainly tests code-generation on small programming problems; it should not be treated as proof of real-world software-engineering superiority.
  • SWE-bench Verified results may overstate capability because of contamination, benchmark artifacts, and the gap between passing tests and producing maintainer-acceptable code.

The Receipts

  1. Claude 4

    Anthropic

  2. Claude 3.5 Sonnet

    Anthropic

  3. SWE-bench Verified Leaderboard

    SWE-bench

  4. SWE-bench

    GitHub Pages (swe-bench.github.io)

  5. SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

    arXiv

  6. SWE-bench Verified: Evaluating Real-World Software Engineering with Grounded and Reliable Tasks

    arXiv

  7. Examining Coding Performance Mismatch on HumanEval and NaturalCodeBench

    arXiv

  8. BigCodeBench: The Next Generation of HumanEval

    Hugging Face / BigCodeBench

  9. SWE-bench Verified

    Vals AI

  10. SWE-Bench Pro (Public Dataset) - Scale Labs

    Scale AI Labs

+ 21 more sources — see the full list on Lenz

Filed Under

HumanEval+SWE-bench Verified

More Fact Checks