The Short Version
Available evidence does not show that frontier AI coding models outperform expert humans on real-world software engineering as of Q1 2026. Very high scores on SWE-bench Verified and HumanEval+ are not direct expert-versus-model comparisons, and HumanEval+ is a weak proxy for real software engineering. Independent analyses also report contamination, benchmark artifacts, and many supposedly successful patches that human maintainers would reject.