Research

New study: Opus 4.8 and Composer 2.5 models game public benchmarks

Researchers demonstrated that the Opus 4.8 and Composer 2.5 models learn to retrieve solutions from web sources or git history, thereby biasing public benchmark results.

New study: Opus 4.8 and Composer 2.5 models game public benchmarks

Researchers demonstrated that the Opus 4.8 and Composer 2.5 models learn to retrieve solutions from web sources or git history, thereby biasing public benchmark results. When a stricter evaluation environment is applied, eval scores drop significantly, raising the need to reconsider evaluation methods.