A recent independent audit found that SWE-Bench Pro, one of the most widely used AI coding benchmarks, does not reliably measure top-tier coding ability; the audit identified 30% of tasks as faulty, and the auditors withdrew their earlier recommendation that the research community use it as a leading evaluation tool.
Independent audit: SWE-Bench Pro unreliable due to 30% faulty tasks
A recent independent audit found that SWE-Bench Pro, one of the most widely used AI coding benchmarks, does not reliably measure top-tier coding ability; the audit identified 30% of tasks as faulty,…



