Friday, Jul 24, 2026 The claims desk. Receipts included. POWERED BY LENZ
IsThis

TECH

The Claim

As of Q1 2026, frontier AI coding models exceed expert human performance on real-world software engineering tasks, as demonstrated by SWE-bench Verified and HumanEval+ results.

The Short Version

Available evidence does not show that frontier AI coding models outperform expert humans on real-world software engineering as of Q1 2026. Very high scores on SWE-bench Verified and HumanEval+ are not direct expert-versus-model comparisons, and HumanEval+ is a weak proxy for real software engineering. Independent analyses also report contamination, benchmark artifacts, and many supposedly successful patches that human maintainers would reject.

Caveats

  • A high benchmark score is not the same as beating expert humans unless both are measured on the same tasks under the same conditions.
  • HumanEval+ mainly tests code-generation on small programming problems; it should not be treated as proof of real-world software-engineering superiority.
  • SWE-bench Verified results may overstate capability because of contamination, benchmark artifacts, and the gap between passing tests and producing maintainer-acceptable code.

The Receipts

  1. Claude 4

    Anthropic

  2. Claude 3.5 Sonnet

    Anthropic

  3. SWE-bench Verified Leaderboard

    SWE-bench

  4. SWE-bench

    GitHub Pages (swe-bench.github.io)

  5. SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

    arXiv

  6. SWE-bench Verified: Evaluating Real-World Software Engineering with Grounded and Reliable Tasks

    arXiv

  7. Examining Coding Performance Mismatch on HumanEval and NaturalCodeBench

    arXiv

  8. BigCodeBench: The Next Generation of HumanEval

    Hugging Face / BigCodeBench

  9. SWE-bench Verified

    Vals AI

  10. SWE-Bench Pro (Public Dataset) - Scale Labs

    Scale AI Labs

+ 21 more sources — see the full list on Lenz

Filed Under

HumanEval+SWE-bench Verified

More Fact Checks