OpenAI audit finds ~30% of SWE-Bench Pro coding tasks are broken
OpenAI's audit of SWE-Bench Pro found 27.4% to 34.1% of 731 coding tasks broken across two review methods. The company retracts its earlier endorsement, citing implications for its Preparedness Framework.