OpenAI retracts its SWE-Bench Pro recommendation after finding ~30% of tasks broken
In "Separating signal from noise in coding evaluations," OpenAI audited Scale AI's SWE-Bench Pro (731-task public split) — the benchmark it had recommended in February as a replacement for SWE-bench Verified — and formally retracted the rec
On 2026-07-08, the verified AI news record added a significant safety, evals & benchmarks development: In "Separating signal from noise in coding evaluations," OpenAI audited Scale AI's SWE-Bench Pro (731-task public split) — the benchmark it had recommended in February as a replacement for SWE-bench Verified — and formally retracted the recommendation. Agent pipeline flagged 200/731 (27.4%); five human engineers flagged 249 (34.1%); 74% overlap; humans found low-coverage tests as the most common issue (9.4% vs pipeline's 4.1%). Failure modes: overly strict hidden tests, underspecified prompts, low-coverage tests, misleading prompts. Frontier scores had inflated from 23.3% to 80.3% in eight months. OpenAI called for new community-built benchmarks rather than naming a replacement.
Context
In "Separating signal from noise in coding evaluations," OpenAI audited Scale AI's SWE-Bench Pro (731-task public split) — the benchmark it had recommended in February as a replacement for SWE-bench Verified — and formally retracted the recommendation. Agent pipeline flagged 200/731 (27.4%); five human engineers flagged 249 (34.1%); 74% overlap; humans found low-coverage tests as the most common issue (9.4% vs pipeline's 4.1%). Failure modes: overly strict hidden tests, underspecified prompts, low-coverage tests, misleading prompts. Frontier scores had inflated from 23.3% to 80.3% in eight months. OpenAI called for new community-built benchmarks rather than naming a replacement. Press: The Stack (2026-07-09), investing.com (2026-07-08), faros.ai audit analysis (2026-07-16). OpenAI's stated motivation ties benchmark validity directly to Preparedness Framework safety cases. Counterargument (from OpenAI itself): agent-pipeline labeling was conservative vs humans, so 30% may understate breakage.
What changed
~30% of SWE-Bench Pro tasks are broken; OpenAI retracts its endorsement. According to OpenAI (PRIMARY, fetched), the supporting record states: “Our datapoint analysis pipeline flagged 200 (27.4%) broken tasks, while the human annotation campaign identified 249 (34.1%)… Given the issues uncovered in this analysis, we retract our earlier recommendation to adopt SWE-Bench Pro." And: "When evaluations have flaws that affect results, they can give a false understanding of capabilities, misrepresenting safety cases and affecting research priorities.”.
Why it matters
Undermines headline coding-benchmark claims across the whole AI-coding market; second OpenAI benchmark reversal in five months; proof that agent-assisted auditing can scale data-quality checks. The safety angle matters because evaluation quality, disclosure, and monitoring determine whether capability claims can be trusted.
Details
The research file records the item under “OpenAI retracts its SWE-Bench Pro recommendation after finding ~30% of tasks broken” with source timing of 2026-07-08. The captured research confidence note is: High (primary). Press: The Stack (2026-07-09), investing.com (2026-07-08), faros.ai audit analysis (2026-07-16). OpenAI's stated motivation ties benchmark validity directly to Preparedness Framework safety cases. Counterargument (from OpenAI itself): agent-pipeline labeling was conservative vs humans, so 30% may understate breakage.
Limitations and caveats
The research file did not identify a blocking caveat, but vendor-supplied claims should still be read as company statements unless independently confirmed.
Sources
Update note: Last reviewed 2026-07-22. Next checkpoint: monitor official channels and the linked source record.
Sources
- OpenAI (PRIMARY, fetched) — official
Drafted with AI assistance from source briefs; reviewed for citation completeness and label accuracy.