The article says benchmark scores reflect not only the model but also the runtime framework and settings, so comparisons require cautious interpretation. It also states that for long-running agents, retaining inference and compressing context allows the model to build on prior knowledge, which can lead to better performance.
AI-generated text
Benchmark scores are affected by the environment and settings as well as the model; retaining inference benefits long-term agents
The article says benchmark scores reflect not only the model but also the runtime framework and settings, so comparisons require cautious interpretation.



