OpenAI audit finds ~30% of SWE-Bench Pro coding tasks are broken
OpenAI's audit of SWE-Bench Pro found 27.4% to 34.1% of 731 coding tasks broken across two review methods. The company retracts its earlier endorsement, citing implications for its Preparedness Framework.
Updated
Why it matters
- Automated pipeline flagged 200 of 731 SWE-Bench Pro tasks (27.4%) as broken
- Human annotation campaign identified 249 broken tasks (34.1%) across the same benchmark
- Frontier models improved from 23.3% to 80.3% pass rate on the 731-task public split in eight months
- OpenAI retracts its earlier recommendation to adopt SWE-Bench Pro
- Four failure categories dominated: overly strict tests, underspecified prompts, low-coverage tests, and misleading prompts
OpenAI's audit found roughly 30% of SWE-Bench Pro tasks are broken. The company's automated pipeline flagged 200 of 731 tasks (27.4%) as defective, and a human annotation campaign with five engineers per task identified 249 broken tasks (34.1%).
The findings, published Tuesday by OpenAI, push the company to retract its earlier endorsement of the benchmark. Scale AI introduced SWE-Bench Pro to replace SWE-bench Verified, the contamination-plagued coding eval that OpenAI publicly flagged in July. The reversal exposes how unstable the ground truth has become for measuring AI coding agents at a moment when labs compete to claim frontier performance on agentic software work.
"SWE-Bench Pro was designed to improve on SWE-bench Verified by testing models on longer horizons and more realistic coding tasks," OpenAI wrote in the post. The 731-task public split tracked rapid progress: frontier models climbed from a 23.3% pass rate to 80.3% in eight months. Those headline numbers now sit on shaky footing.
What did OpenAI actually audit?
OpenAI ran a two-track review of flagged SWE-Bench Pro tasks. The first track used a datapoint analysis pipeline that scanned task instructions, model attempts, and the tests used to grade them. The pipeline flagged 286 potentially broken tasks out of 731.
The second track applied two independent audits to that 286-task subset:
- A human-supervised agent review using Codex-based investigator agents with access to the task repository and environment. After several repeats, a researcher made a final judgment.
- A human annotation campaign with five experienced software engineers per task. Reviewers formed an independent view from the visible problem statement, test cases, and the gold patch before consulting agent analysis.
The reviewers had no stake in the original benchmark. OpenAI trained them on benchmark goals, issue taxonomy, and edge cases before the review began.
What kinds of broken did they find?
Four failure categories dominated the audit:
- Overly strict tests enforce implementation details the prompt never specified, invalidating correct solutions.
- Underspecified prompts omit requirements that hidden tests later demand.
- Low-coverage tests under-check the requested feature, so incomplete fixes pass.
- Misleading prompts steer models toward the wrong behavior or contradict what tests require.
The mix matters. OpenAI's agent pipeline and human reviewers agreed on category in 74% of flagged cases. Humans, however, were likelier to assign multiple labels to a single task, suggesting layered problems rather than clean single failures.
The largest gap between the two methods appeared on low-coverage tests. Humans marked 9.4% of the benchmark as having low-coverage tests as the most common issue. The agent pipeline landed on 4.1%. Humans caught what the agents missed.
Why does this matter beyond one benchmark?
SWE-Bench Pro was meant to be the clean replacement. SWE-bench Verified, the prior industry standard, suffered from contamination and design issues that OpenAI flagged publicly in July. Scale AI's Pro version added longer horizons and tasks drawn from both public and private repositories.
The Pro audit now exposes a second-order problem. Tasks pulled from real GitHub histories inherit the messiness of human collaboration. Pull request tests often validate a specific change rather than an implementation-agnostic standard. Issue descriptions, merged code, and unit tests do not always line up into clean, isolated evaluation items.
"Issues and pull requests from open-source repositories were originally created for human collaboration, often through long back-and-forths between maintainers and contributors," OpenAI wrote. "As a result, problem descriptions, merged code, and unit tests do not always line up to form clean, isolated tasks for evaluating models reliably."
For labs, the practical risk is concrete. A 30% broken rate inflates measured capability and can misroute safety cases. OpenAI ties benchmark validity directly to its Preparedness Framework, the document guiding the company's deployment and safety decisions.
How did the two review methods compare?
Human reviewers were more skeptical than investigator agents. In no flagged task did "not broken" emerge as the most common human label. The agent pipeline undercounted cases where reviewers saw additional or overlapping issues, particularly on low-coverage tests.
The pattern points to a useful asymmetry. Agents scale. Humans catch edge cases the agents miss. OpenAI's takeaway: combine both, and treat the resulting label as a conservative lower bound on the true broken rate.
That conservatism shapes the headline number. OpenAI's "~30% of SWE-Bench Pro tasks are broken" estimate sits between the pipeline's 27.4% and the human campaign's 34.1%. The company frames it as a floor, not a ceiling.
The audit also surfaces a quiet benefit of stronger models. As capabilities improve, labs can use the same models to inspect prompts, tests, patches, traces, and edge cases with greater depth. Codex-based investigator agents already run the first pass in OpenAI's pipeline. Future audits could fold human-style review into an agent loop, cutting the cost of catching broken tasks at scale.
What happens now?
OpenAI is retracting its recommendation to adopt SWE-Bench Pro. The reversal comes eight months after the same lab pushed the community toward Pro as the better instrument for tracking agentic coding.
The post calls for new benchmarks built by experienced software developers specifically to test model capability, with human oversight baked into curation. "We hope the wider evaluation community will develop new benchmarks built by experienced software developers specifically to test model capabilities," OpenAI wrote.
Frontier coding scores now sit on benchmarks that two independent audit methods found roughly one-third defective. Labs reporting 80% pass rates on Pro will need to reconcile those numbers with the broken-task count, or find a better instrument.
"Ultimately, an eval should provide meaningful signal through benchmarks that are hard to game, easy to trust, and genuinely reflective of model capability or alignment," OpenAI wrote. "Because these results inform OpenAI's deployment and safety decisions, the evals we track need to be valid and informative."
The audit lands at a moment when coding benchmarks drive procurement, investor narratives, and safety cases. A 30% broken rate is not a rounding error. It is a measurement problem the industry will need to fix before the next frontier release.
Original: cdn.openai.com
More from Rebecca Stone
Show full bio
Correspondent covering consumer brands and retail at AI In Context.
214 articles