Source: ai-research/anthropic-claude-sonnet-5-5-announcement-2026-09-28.md — Anthropic’s launch page (anthropic.com/claude-sonnet-5-5, dated 2026-09-28), fetched 2026-09-29. Launch post also mirrored on r/ClaudeAI (raw/reddit-1wslxzs.md, score 2,115). Claude Code facts from v2.1.284; SDK facts from anthropic-sdk-python v1.9.0 and TypeScript sdk-v0.129.0.
Claude Sonnet 5.5 (claude-sonnet-5-5) shipped 2026-09-28, six days after Opus 5.5, as the second model in the Claude 5.5 family. It keeps Sonnet 5’s price (10) but needs far fewer tokens, runs 30%+ faster, and on several benchmarks lands within a point or two of Opus 5.5 at half the per-token price. Anthropic positions it as “strongest at well-scoped everyday tasks, fixing bugs, and creating polished documents, slides, and spreadsheets,” and keeps Opus 5.5 for “complex work requiring careful judgment.”
Key Takeaways
- Same price, fewer tokens, faster. 10 output / **2.50 cache write per MTok, the same as Sonnet 5. Anthropic reports up to 30% lower cost per task and 30%+ faster output, its fastest Sonnet to date. Opus 5.5 costs exactly 2× per input and output token, with the same $0.20 cache-read rate.
- Close to Opus 5.5 on paper. GDPval-AA v2.1 1844 vs 1846; CursorBench 4.0 55.5% vs 57.8%; OSWorld 2.1 80.1% vs 81.8%. It beats Opus 5.5 on Terminal-Bench 4.0 (70.6% vs 66.4%). Anthropic still says “Opus 5.5 remains clearly stronger at complex, open-ended work requiring sustained judgment.”
- Cheap at low effort. On several benchmarks, Low or Medium effort beats Sonnet 5’s best score for about a tenth of the cost per task. Default effort: Medium in Claude Code and the Claude apps, High on the Claude Platform.
- Max effort can score lower than Xhigh. On FrontierCode, Sonnet 5.5 scored 46.2% at Max vs 52.1% at Xhigh. At Max it more often ran Claude Code’s code-review skill across many subagents, which led to timeouts or out-of-scope edits that the benchmark penalizes. More effort is not automatically better on tightly scoped tasks — see Model vs. Effort.
- A planner/implementer split is now a vendor-quoted pattern. One tester: “When Claude Opus 5.5 sets the architecture and general framework for a game, I would feel confident in letting Sonnet 5.5 implement it.” See Two-Model Workflow for the mechanics and limits of running that split in one session.
- First Sonnet with cyber safeguards and anti-extraction classifiers. Its cyber capability is “comparable to Opus 5’s,” so higher-risk cybersecurity tasks visibly fall back to Sonnet 5; routine bug-fixing is unaffected. Biology safeguards are unchanged from Sonnet 5. It also expands preserved thinking so “Claude’s thinking cannot be decoupled from the account that created it,” which affects anyone who moves conversations between accounts, including switching accounts mid-session in Claude Code.
- Migration gotcha: thinking-off users must switch settings. “If you run Sonnet with thinking off, you’ll need to switch to the new
between_toolssetting, which keeps up-front thinking off, before moving to Sonnet 5.5.” Thebetween_toolsthinking type shipped in Python SDK v1.9.0 and TypeScript sdk-v0.129.0. - Zero data retention available, as with Opus 5.5 and Sonnet 5. Haiku 5.5 “will join the Claude 5.5 family in the coming weeks.”
Benchmarks (vendor-reported)
| Benchmark | Sonnet 5.5 | Sonnet 5 | Opus 5.5 | GPT-6 Sol |
|---|---|---|---|---|
| Terminal-Bench 4.0 | 70.6% | 10.3% | 66.4% (xhigh) | — |
| FrontierCode 1.1 Main | 46.2% Max / 52.1% Xhigh | 42.4% | 54.4% | 49.3% |
| CursorBench 4.0 | 55.5% | 34.1% | 57.8% | — |
| GDPval-AA v2.1 (Elo) | 1844 | 1449 | 1846 | 1487 |
| AA-Briefcase v1.1 (Elo) | 1811 | 1359 | 1822 | 1483 |
| Humanity’s Last Exam (tools) | 64.5% | 54.9% | 67.7% | — |
| OSWorld 2.1 (partial) | 80.1% | 57.0% | 81.8% | — |
| Chartography (no tools) | 61.6% | 15.6% | 64.4% | 53.6% |
- Watch the benchmark version. Sonnet 5’s 10.3% is on Terminal-Bench 4.0. Older wiki figures for Sonnet 5 on earlier Terminal-Bench versions are not comparable; this is a harder benchmark, not a Sonnet 5 regression.
- Artificial Analysis ran GDPval-AA and AA-Briefcase on a pre-release deployment with a structured-outputs bug (since fixed). Anthropic expects the effect “if any, to be small.” GPT-6 Sol’s AA and Chartography scores may predate an OpenAI image-understanding fix.
What testers reported (vendor-selected)
- Fewer iterations. Base44: across 118 real app builds, apps “scored level with Opus 5” in 3.6 iterations per build vs 7.7, with the fewest failed tool calls of any model compared, and it rarely stopped mid-build to ask a question.
- Token use on finance work. On 2,441 finance tasks, it scored ahead of Sonnet 5 using ~121k tokens per answer vs 497k.
- Support and chat. Zendesk: tickets processed 20% faster, with fewer wrong decisions. Slack: better on “almost all” offline Slackbot evals, with ~14% fewer output tokens and no prompt changes.
- Decks. Given earnings materials, call transcripts, and a slide template, its 10-slide operating-review first draft was judged “ready to send as is” by two experts in Anthropic’s internal test.
Field reports (community, unverified)
- Sonnet 5.5 vs Opus 5.5 on one prompt (
raw/reddit-1wspepd.md, r/ClaudeCode, score 142). The prompt: one-shot a Three.js cartoon planet with orbit controls. The first run had both models on one machine sharing CPU and GPU, which the author says contaminated the timings: Sonnet 5.5 took 10.8 min and cost 4.79. The sequential re-run is the usable number: Opus 5.5 14.0 min, 56.6k output tokens, 1.69 (range 7.0–15.2 min, 2.11). The author’s summary, “~20% of the cost,” matches the contaminated first run; the sequential runs put Sonnet at about two-thirds of Opus’s cost. ^[inferred] He made Sonnet 5.5 his default builder. - Where multi-agent cost actually goes (
raw/reddit-1wsx03y.md, r/ClaudeCode, score 308). Six Sonnet 5.5 agents at High effort (1 lead, 5 builders) one-shot a Three.js Mario Kart clone in 67 minutes: 424 model calls, 889k tokens written (352k of them thinking), 78M tokens read, of which 76M were agents re-reading context they had already loaded, for **0.20) dominates the bill, not the output price. ^[inferred] See Prompt Caching for Agencies. - Hard to tell from Opus 5.5 in daily use (
raw/Sonnet_5.5_Is_Here._Look_What_It_Can_Build..md, Matthew Berman, after several days of use): “I can’t really tell the difference between Sonnet 5.5 and Opus 5.5. Maybe I’m not giving it hard enough tests.” Two quirks: it writes British spellings (“colour”) by default, which a writing rule fixes, and its generated game sound and music “really doesn’t sound good at all.”
Claude Code and SDK changes that shipped with it
- Claude Code v2.1.284 (2026-09-28):
claude-sonnet-5-5added and made the default Sonnet model on the Anthropic API (1M context, 10, $0.20 cache reads). The release note scopes the default to the Anthropic API; defaults on Bedrock, Vertex, and Foundry are not stated. The same release starts interactive sessions in auto mode when no permission mode is configured. See Week 40. - SDKs (2026-09-28):
claude-sonnet-5-5in Python v1.9.0, TypeScript sdk-v0.129.0, and the Bedrock, Vertex, and Foundry SDKs, with thebetween_toolsthinking type and cache diagnostics now GA. See SDK releases, Aug–Sep 2026.
Implementation
Tool/Service: Claude Sonnet 5.5 (claude-sonnet-5-5)
Setup: The Claude apps, Claude Code, the Claude Platform, and Amazon Web Services, Google Cloud, and Microsoft Azure (“available on all platforms”).
Cost: 10 output / 2.50 cache write per MTok.
Integration notes: Move thinking-off configurations to between_tools before switching. Sessions that change accounts mid-conversation will lose the ability to reuse prior thinking.
Try It
- Route implementation to Sonnet 5.5, planning to Opus 5.5. Test the split the launch quote describes: Opus writes the plan, Sonnet builds it, and compare the combined cost with Opus alone.
- Test Low and Medium effort first. Anthropic’s charts say Low/Medium beats Sonnet 5’s best on several benchmarks at about a tenth of the cost.
- Don’t assume Max is best. For scoped tasks such as a single-file fix, compare Xhigh and Max. The FrontierCode footnote shows Max doing out-of-scope work.
- If you were on Sonnet 5 for cost, switch now. Same price, fewer tokens per task.
Open Questions
- System card not read. A Sonnet 5.5 system card (2026-09-28) exists; this article uses the launch page only.
- Default on Bedrock / Vertex / Foundry in Claude Code is not stated in v2.1.284.
- Independent benchmarks. All numbers above are Anthropic’s or its partners’. Community reports from the first day were mostly single-prompt demos (games, animations), not controlled comparisons.
- Haiku 5.5 release date not given beyond “the coming weeks.”
Related
- Claude Opus 5.5 — the first 5.5 model; 2× the per-token price, stronger on open-ended work.
- Claude Sonnet 5 — the predecessor at the same price.
- Two-Model Workflow — how a planner/implementer split actually runs.
- Model vs. Effort — the effort dial, including why Max can underperform.
- Cost & Intelligence Levers — where a half-price near-peer model changes routing.
- Week 40 digest — the Claude Code release that made it the default Sonnet.