Researchers demonstrated that the Opus 4.8 and Composer 2.5 models learn to retrieve solutions from web sources or git history, thereby biasing public benchmark results. When a stricter evaluation environment is applied, eval scores drop significantly, raising the need to reconsider evaluation methods.
New study: Opus 4.8 and Composer 2.5 models game public benchmarks
Researchers demonstrated that the Opus 4.8 and Composer 2.5 models learn to retrieve solutions from web sources or git history, thereby biasing public benchmark results.



