Research

Fixed Compute Budgets Skew AI Benchmark Results, UK Study Finds

The UK's AI Security Institute tested frontier models on seven benchmarks and found that fixed compute budgets baked into evaluations systematically understate model capabilities.

Fixed Compute Budgets Skew AI Benchmark Results, UK Study Finds

The UK’s AI Security Institute (AISI) tested frontier AI models across seven benchmarks and concluded that widely relied-on scores are systematically flawed. The common issue is a fixed compute budget embedded in every test: models are allowed only a defined amount of tokens or computational steps to produce answers.

Findings from the AISI tests

  • When agents are given more tokens, their performance keeps improving: up to 25% higher on software tasks and 22% higher on math tasks.
  • A cybersecurity challenge that would take a human expert about 20 hours remained unsolved under standard budget cutoffs. The task was only solved once models were allowed to run up to 30 million tokens, well beyond typical benchmark limits.
  • Removing the cap also changed measured progress rates: AISI’s own cyber progress rate shifted from doubling every 67 days to doubling every 40 days once the budget constraint was lifted.

Why this matters in practice

AISI’s results show that capability is a curve that rises with computational spend. Most leaderboards, headlines claiming “model X scores Y,” and safety evaluations relied on by regulators are therefore measuring performance under an arbitrary budget, not intrinsic capability. Those scores are not conservative estimates — they are budget-dependent and can be misleading when presented without context.

Implications for rankings and regulation

The institute warns that every number used to rank, market, or regulate models is effectively an artifact of a chosen compute budget rather than a standalone measure of intelligence. To address this, benchmarking practices should either adopt more flexible measurement frameworks or explicitly document how budget constraints affect results so that scores can be interpreted correctly.

Summary

The AISI study demonstrates that fixed compute caps cause many benchmarks to understate what models can do. Until benchmarking methods account for token- and compute-dependence, published scores and the conclusions drawn from them risk being misleading.