Model launches

Anthropic launches Claude Opus 4.8 with faster, more reliable and resource‑controllable enterprise AI

Anthropic released Claude Opus 4.8 on May 28, 2026, upgrading Opus 4.7 with improved benchmark performance, greater reliability in agentic tasks, and new user controls for computational effort.

Anthropic launches Claude Opus 4.8 with faster, more reliable and resource‑controllable enterprise AI

Anthropic released Claude Opus 4.8 on May 28, 2026, upgrading Opus 4.7 with improvements across benchmarks, agentic reliability, and new user controls for resource use. The company says the new model produces better results on coding, agentic skills, reasoning, and practical knowledge tasks while introducing operational features aimed at enterprise workflows.

Capabilities of Opus 4.8

Anthropic compares Opus 4.8 with its predecessor and other models across various tests. Key points highlighted by the company include:

  • improved judgment and reliability on agentic tasks;
  • a fast mode that runs at 2.5× the speed of normal operation;
  • better handling of long-session context and style direction;
  • reduced likelihood of letting code flaws pass unnoticed and more frequent uncertainty flagging.

Anthropic reports Opus 4.8 is about four times less likely than Opus 4.7 to allow flaws in code it wrote to pass without remark. Detailed capability evaluations and the alignment assessment are published in the Claude Opus 4.8 System Card.

Early user feedback and benchmark highlights

Early testers quoted by Anthropic say Opus 4.8 asks better questions in Claude Code, detects its own mistakes, pushes back on unsound plans, and sustains complex multi-service explorations. Several customers and partners — including Tom Pritchard (Staff Engineer), Kay Zhu (Co-Founder and CTO), and Michael Truell (Co-Founder and CEO) — reported improved end-to-end performance on multi-step and tool-using tasks.

Notable benchmark outcomes cited in the announcement:

  • On the Super-Agent benchmark, Opus 4.8 is reported as the only model to complete every case end-to-end, outperforming prior Opus models and GPT-5.5 at cost parity.
  • Opus 4.8 scored 84% on Online-Mind2Web, a meaningful jump over Opus 4.7 and GPT-5.5.
  • On the Legal Agent Benchmark, Opus 4.8 recorded the highest score to date and was the first model to exceed a 10% all-pass standard overall.

Enterprise partners such as Databricks (Genie) and Hebbia reported improvements in agentic reasoning, citation precision, and token efficiency for document- and finance-focused workflows.

Honesty and alignment

A prominent improvement Anthropic calls out is the model’s increased “honesty”: Opus 4.8 more often flags uncertainty and makes fewer unsupported claims. The Alignment team concluded Opus 4.8 "reaches new highs on our measures of prosocial traits like supporting user autonomy and acting in the user’s best interest." Rates of misaligned behavior (e.g., deception or cooperation with misuse) are substantially lower than Opus 4.7 and similar to Anthropic’s best-aligned model, Claude Mythos Preview. Full alignment and pre-deployment safety test results are available in the System Card.

New features: dynamic workflows and effort control

In addition to the model itself, Anthropic is shipping several features:

  • Dynamic workflows (research preview): In Claude Code, Claude can plan large tasks, spawn hundreds of parallel subagents within a single session (with Opus 4.8 these agents can run longer), and verify outputs before reporting back. Anthropic gives the example of performing codebase-scale migrations across hundreds of thousands of lines from kickoff to merge using the existing test suite as the pass criterion. Dynamic workflows will be available for Enterprise, Team, and Max Claude Code plans.
  • Effort control in claude.ai and Cowork: A new control next to the model selector lets users choose how much effort Claude invests in a response. Higher effort settings produce more frequent and deeper thinking for better responses; lower effort settings yield faster replies and slower consumption of rate limits. This control is available on all plans.
  • Messages API update: The Messages API now accepts system entries inside the messages array, allowing developers to update Claude’s instructions mid-task without breaking the prompt cache or routing the update through a user turn. This supports live changes to permissions, token budgets, or environment context while an agent runs.

Effort defaults and token use

Opus 4.8 defaults to high effort, which Anthropic says balances quality and user experience. For coding tasks, this default spends a similar number of tokens as Opus 4.7’s default but achieves better performance. Users may select "extra" ("xhigh" in Claude Code) or "max" to expend more tokens for improved results; Anthropic recommends "extra" for difficult tasks and long-running asynchronous workflows. Rate limits in Claude Code have been increased to accommodate higher token usage at higher effort levels.

Pricing and availability

Claude Opus 4.8 is available worldwide as of May 28, 2026. Regular pricing remains unchanged from Opus 4.7: $5 per million input tokens and $25 per million output tokens. Fast mode pricing is $10 per million input tokens and $50 per million output tokens; Anthropic states the fast mode now runs 2.5× faster and is three times cheaper than prior-model fast modes. Developers can access the model named "claude-opus-4-8" via the Claude API.

What’s next

Anthropic calls Opus 4.8 a modest but tangible improvement over its predecessor and says it is working on lower‑cost models offering similar capabilities. The company is also developing a higher‑capability class of models under Project Glasswing; a small set of organizations currently use Claude Mythos Preview for cybersecurity tasks. Anthropic says stronger cybersecurity safeguards are required before wider release of Mythos‑class models and that it is rapidly developing those protections with the aim of broader availability in the coming weeks.

Notes

The announcement references changes to benchmark harnesses and scoring (Terminal-Bench 2.1, OSWorld-Verified) and provides further methodological detail in the Claude Opus 4.8 System Card. Overall, Anthropic positions Opus 4.8 as a step forward for enterprise agentic workflows, codebase tasks, and high-stakes professional use cases.