Can We Trust AI Benchmarks Anymore? What OpenAI's SWE-Bench Audit Means for You
OpenAI audited the SWE-Bench Pro coding benchmark and found ~30% of tasks are broken. Every model leaderboard based on it is now suspect. Here's what that means for choosing AI coding tools — and why the future of evaluation is AI auditing AI.