OpenAI no longer recommends SWE-Bench Pro

OpenAI has withdrawn its recommendation for the SWE-Bench Pro coding benchmark after an audit revealed that approximately 30% of its tasks are broken. The company emphasizes the need for accurate evaluation metrics to properly assess AI model capabilities and safety.
Why it matters
Reliable benchmarks are critical for the AI industry to track progress and ensure that safety claims are based on valid data.
Through a detailed audit, we find widespread task issues in SWE-Bench Pro and estimate that ~30% of the tasks are broken.
The report focuses on technical transparency and the integrity of AI evaluation standards.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in