The Sequence AI of the Week #895: OpenAI's Show Us Where Coding Evals Break
A visual explanation for OpenAI's new science for coding evaluations and benchmarks.

TL;DR
- OpenAI conducted an audit of the SWE-Bench Pro coding benchmark.
- The audit found that roughly 30% of the benchmark tasks are defective.
- Both AI agents and human software engineers identified flawed tasks.
- OpenAI has withdrawn its recommendation for the field to adopt SWE-Bench Pro.
- The findings suggest that precise scores on benchmarks may not accurately reflect true coding ability.