Model launches

Anthropic releases Claude Sonnet 5: a more agentic, lower‑cost midsize model

Anthropic has introduced Claude Sonnet 5, a midsize model with improved agentic abilities that can plan, use tools and run more autonomously than previous versions.

Anthropic releases Claude Sonnet 5: a more agentic, lower‑cost midsize model

Anthropic has unveiled Claude Sonnet 5, a more powerful, agentic iteration of its midsize model. According to the company, Sonnet 5 can plan, use tools like browsers and terminals, and operate autonomously at levels that until recently required larger, costlier models.

What changed and why it matters

The release underscores that agentic capability is now a baseline expectation across price tiers: the competition is increasingly about who can provide agentic behavior most cheaply and reliably without human oversight.

Availability and pricing

Claude Sonnet 5 became the default model for the Free and Pro plans starting Tuesday and is available to all subscriptions. Introductory pricing runs through August 31 at $2 per million input tokens and $10 per million output tokens; after August 31 the input price rises to $3 per million while output remains $10 per million. Anthropic says these rates make Sonnet 5 cheaper than Opus 4.8 as well as OpenAI’s GPT‑5.5 and Google’s Gemini 3.1 Pro, though it remains more expensive than Gemini 3.5 Flash.

Performance and benchmarks

Anthropic reports materially better agentic performance versus Sonnet 4.6 (released in February) across reasoning, tool use, software coding and knowledge work.

  • On an agentic coding benchmark Sonnet 5 scores 63.2%, compared with Opus 4.8 at 69.2% and Sonnet 4.6 at 58.1%.
  • On a knowledge‑work benchmark Sonnet 5 slightly outperformed Opus 4.8, which is known for handling the hardest problems such as subtle judgment calls and deep research. Anthropic notes that while Opus 4.8 remains the choice for higher accuracy on the most demanding tasks, Sonnet 5 offers developers lower‑cost options that are a big quality step up from previous Sonnet versions.

User reports

Testers cited in Anthropic’s blog post say Sonnet 5 is better at finishing complex workflows that earlier versions would stall on, and that it often checks its own outputs without being explicitly prompted. Daniel Shepard, senior engineer at Zapier, said the model completed a two‑part job end‑to‑end (updating Salesforce account tiers and sending a launch announcement to enterprise contacts) where earlier models had stalled. Shepard called it "a no‑brainer" for day‑to‑day automation.

Safety and limitations

According to Anthropic, Sonnet 5 exhibits lower rates of undesirable behaviors—such as cooperating with misuse, deception, hallucination and sycophantic responses—than Sonnet 4.6. It also more reliably refuses malicious requests and resists hijack attempts in prompt‑injection attacks. However, the company acknowledges Sonnet 5 does not match Opus 4.8 or Claude Mythos Preview in terms of alignment for the most safety‑critical tasks: evaluations show it has much lower ability to perform dangerous cybersecurity tasks than current Opus models. Fabian Hedin, co‑founder of Lovable, said Claude Sonnet 5 "refuses unsafe requests cleanly and consistently," adding that tools that know when to say no are as important as tools that know how to build.

Conclusion

Claude Sonnet 5 is Anthropic’s push to make agentic functionality affordable at the midsize tier. It narrows the gap between smaller, cheaper models and larger, more capable ones by improving autonomous behavior and safety over its predecessor, while Anthropic keeps Opus models positioned for the highest‑accuracy and most safety‑sensitive use cases.