Model launches

Anthropic releases Claude Opus 4.8 with bugfix-focused improvements, not a major leap

Anthropic launched Claude Opus 4.8 forty-two days after Opus 4.7, the shortest interval between releases in the company’s history.

Anthropic releases Claude Opus 4.8 with bugfix-focused improvements, not a major leap

Anthropic has released Claude Opus 4.8, forty-two days after Opus 4.7 — the shortest interval between releases in the company’s history. The update preserves the same context window and keeps pricing unchanged at $5 and $25 per million tokens.

What changed

  • Most independent benchmarks show score increases compared with Opus 4.7. The notable exception is the agentic coding evaluation Terminal-Bench 2.1, where GPT-5.5 continues to lead.
  • Opus 4.6 has been removed from the web interface.
  • According to the release notes and user reports, Opus 4.8 primarily addresses prior issues: it reduces hallucinations, tightens instruction-following, and remedies certain performance problems described by users as “laziness.”

Why this matters

Since Opus 4.5, Anthropic has operated under an implicit expectation that each new generation will advance on multiple axes. Two recent releases failed to meet that expectation: Opus 4.7 shipped in a state that many users found rough, prompting some to turn temporarily to alternatives such as Codex. Opus 4.8 responds to those deficiencies, but the changes largely read as fixes rather than transformative feature additions.

Notably, creative output quality in Opus 4.8 is reported to be worse than in Opus 4.6, meaning progress has not been uniform across capabilities. Meanwhile, in at least one important coding benchmark (Terminal-Bench 2.1), a competing model, GPT-5.5, remains ahead.

Conclusion

The Opus 4.8 release appears to be focused on repairing regressions and improving reliability rather than delivering a major step forward. While several benchmarks improved and some persistent issues were addressed, the removal of older versions from the interface and uneven performance across tasks leave questions about Anthropic’s release strategy and product roadmap. Future releases will need to demonstrate both stability and broad-based performance gains to restore confidence among users and evaluators.