Claude Sonnet 5.5 is out: 30% faster, same $2/$10 price, and a Terminal-Bench score that beats Opus 5.5
Anthropic's second Claude 5.5 model posts a 70.6% Terminal-Bench coding score at unchanged pricing, though independent testing shows the per-task savings depend on the effort setting.
Anthropic released Claude Sonnet 5.5 on Monday, the second model in its Claude 5.5 family after last week's Opus 5.5 launch. The headline numbers: output arrives more than 30% faster than Sonnet 5, and Anthropic says it costs up to 30% less per task, at the same list price of $2 per million input tokens and $10 per million output tokens.
The coding benchmarks are the real story. On Terminal-Bench 4.0, an agentic coding evaluation, Sonnet 5.5 scored 70.6%, up from Sonnet 5's 10.3% and ahead of Opus 5.5's 66.4%. On CursorBench 4.0, which replays tasks from real Cursor coding sessions, it hit 55.5% against Sonnet 5's 34.1%, landing within two points of Opus 5.5. On FrontierCode 1.1 at High effort, it scored 10 points above Sonnet 5 at the same setting for about one fifteenth the cost per task, by Anthropic's accounting. It is also the first Sonnet to beat Pokémon Red working only from screenshots.
The pricing math: $2/$10 per million tokens in and out, $0.20 per million cache reads, $2.50 for cache writes, all unchanged from Sonnet 5. The savings come from using fewer tokens and batching tool calls, not from cheaper tokens. Early testers backed that up. Epic Games COO Daniel Vogel said Sonnet 5.5 "cleared the same quality bar you'd expect from a higher-tier model" in a system design audit and data flow review, managing tens of thousands of lines of gameplay code. Slack reported about 14% fewer output tokens on its Slackbot evals, Box saw runs 2.4x faster, and Zendesk processed tickets 20% faster. Balyasny Asset Management ran 2,441 finance tasks and watched tokens per answer fall from roughly 497,000 to about 121,000.
Two things to keep in mind. First, Sonnet 5.5 is the first Sonnet to ship with the cyber safeguards and anti-distillation classifiers Anthropic previously reserved for its biggest models. Higher-risk cybersecurity tasks will visibly fall back to Sonnet 5, and the model's thinking is now bound to the account that created it, so moving conversations between accounts, including switching accounts mid-session in Claude Code, works differently.
Second, the "up to 30% less per task" claim carries an asterisk. Independent testing by Artificial Analysis found that at Max effort, Sonnet 5.5 used about 193,000 output tokens per task, the heaviest token use they have measured, roughly 60% more than Opus 5.5 at max and about seven times GPT-6 Astra, pushing its cost per task to $7.60, around 50% higher than Sonnet 5's. Anthropic's own charts show the savings at lower effort levels, where Sonnet 5.5 beats Sonnet 5's best score for about a tenth of the cost per task. Read it as: cheaper for everyday work at default settings, pricier when you crank the effort dial.
Anthropic is careful to say Opus 5.5 remains the stronger model for complex, open-ended work requiring sustained judgment, and that benchmark scores capture only one facet of capability. Sonnet 5.5 is available now on the Claude Platform as claude-sonnet-5-5, plus AWS, Google Cloud and Azure, with zero data retention. Haiku 5.5, the small fast sibling, arrives in the coming weeks. The release lands as Anthropic expands its lineup ahead of a planned IPO, with enterprise customers accounting for about 80% of its business.