The UK’s AI Security Institute (AISI) tested frontier AI models across seven benchmarks and concluded that widely relied-on scores are systematically flawed. The common issue is a fixed compute budget embedded in every test: models are allowed only a defined amount of tokens or computational steps to produce answers.
Findings from the AISI tests
- When agents are given more tokens, their performance keeps improving: up to 25% higher on software tasks and 22% higher on math tasks.
- A cybersecurity challenge that would take a human expert about 20 hours remained unsolved under standard budget cutoffs. The task was only solved once models were allowed to run up to 30 million tokens, well beyond typical benchmark limits.
- Removing the cap also changed measured progress rates: AISI’s own cyber progress rate shifted from doubling every 67 days to doubling every 40 days once the budget constraint was lifted.
Why this matters in practice
AISI’s results show that capability is a curve that rises with computational spend. Most leaderboards, headlines claiming “model X scores Y,” and safety evaluations relied on by regulators are therefore measuring performance under an arbitrary budget, not intrinsic capability. Those scores are not conservative estimates — they are budget-dependent and can be misleading when presented without context.
Implications for rankings and regulation
The institute warns that every number used to rank, market, or regulate models is effectively an artifact of a chosen compute budget rather than a standalone measure of intelligence. To address this, benchmarking practices should either adopt more flexible measurement frameworks or explicitly document how budget constraints affect results so that scores can be interpreted correctly.
Summary
The AISI study demonstrates that fixed compute caps cause many benchmarks to understate what models can do. Until benchmarking methods account for token- and compute-dependence, published scores and the conclusions drawn from them risk being misleading.



