Artificial Analysis (AA), a widely cited AI benchmark, published a score of 61 for GPT-6 Astra on September 3, placing it below Muse Spark 1.3. Within 24 hours the benchmark index was updated to version 4.2, and in the revised index Astra moved to second place.
Artificial Analysis said the change was a methodology update prepared over several months. According to AA, the adjustment was not an ad-hoc correction but a planned methodological revision.
Why this matters
Benchmarks are expected to provide objective, reproducible comparisons: a given score should serve as a stable point of reference. In this case, the score and the associated ranking changed the day after publication, illustrating that the numbers the industry frequently cites can be fluid.
After the update there were no widely reported retractions or withdrawals of earlier quotes — neither labs nor media outlets publicly challenged which other elements of the methodology had been changed. The episode therefore highlights that benchmark figures used in press releases and marketing can be edited and are not necessarily immutable measurements.
Concrete facts
- Date: September 3 — initial score published.
- Initial result: GPT-6 Astra scored 61, below Muse Spark 1.3.
- Change: index updated to v4.2 within 24 hours.
- AA's statement: the revision was a methodological update prepared over months.
The incident did not produce a public scandal or formal retractions, but it underscores that the most-quoted benchmark scores in AI may be part of indices that can be revised after publication.



