The Sequence AI of the Week #895: OpenAI's Show Us Where Coding Evals Break

A visual explanation for OpenAI's new science for coding evaluations and benchmarks.

The Sequence AI of the Week #895: OpenAI's Show Us Where Coding Evals Break

TL;DR

  • OpenAI conducted an audit of the SWE-Bench Pro coding benchmark.
  • The audit found that roughly 30% of the benchmark tasks are defective.
  • Both AI agents and human software engineers identified flawed tasks.
  • OpenAI has withdrawn its recommendation for the field to adopt SWE-Bench Pro.
  • The findings suggest that precise scores on benchmarks may not accurately reflect true coding ability.