Sonnet 5.5 is built for focused coding, fast iteration, and polished documents, slides, and spreadsheets. Keep Opus 5.5 for the most complex work that demands sustained judgment.
A faster everyday complement to Opus
Anthropic introduced Claude Sonnet 5.5 on September 28, 2026. It supports text and image input, text output, a one-million-token context window, and up to 128,000 output tokens in standard requests. Adaptive reasoning offers five effort levels, from low to max.
Standard API rates are $2 per million input tokens, $10 per million output tokens, and $0.20 per million cache-read tokens. These are unchanged from Sonnet 5. Anthropic’s claimed savings come from using fewer tokens on typical work, not a token-price cut.
At max effort, Sonnet 5.5 scores 56 on Artificial Analysis’s Intelligence Index v4.3.2, close to Opus 5.5’s 58. But max is not always the economical setting: AA estimates $7.60 per Index task for Sonnet versus $5.98 for Opus at max. Match the effort to the job.
Anthropic reports major coding gains and stronger visual understanding, while noting that Opus 5.5 remains better at complex, open-ended work. Model access in SPIRITT depends on the models and connections enabled for your workspace.
1M context128K standard outputText + image inputFive effort levels$2 / $10 API pricingRelease brief
Artificial Analysis Intelligence Index
Index v4.3.2, September 28, 2026. Selected configurations; rounded index points, not percentages. Claude uses adaptive reasoning; Sonnet 5.5, Opus 5.5 and Fable 5.1 use Default Fallback. Sonnet max and Opus xhigh both round to 56; their tiny unrounded gap does not establish a meaningful difference.
Anthropic-reported evaluation
Terminal-Bench 4.0: Anthropic
Anthropic’s Claude Code bare-mode run: 66 tasks, cached dependencies, no internet egress. Sonnet and Opus 5.5 used five trials per task with safeguards; fallback affected 1.2% and 2.5% of requests respectively. Sonnet is max; Opus is xhigh. Other comparators’ repeat counts differ. Separate from AA’s run.
Terminal-Bench 4.0: AA
Artificial Analysis: mini-swe-agent, 66 tasks, three repeats, pass@1, 500-step cap. All rows use max effort; Claude 5.5 configurations and Fable use Default Fallback. These results must not be substituted for Anthropic’s Claude Code run.
FrontierCode v1.1 Main
Cognition’s 150-task evaluation, reported by Anthropic, using Claude Code and Codex CLI. Composite scores, not simple test pass rates. Sonnet scores lower at max than xhigh. In two cases Cognition examined, extra code-review agents caused a timeout or out-of-scope edits.
CursorBench 4.0
Cursor’s production-agent evaluation, reported by Anthropic. Max-effort scores shown. Sonnet’s reported task costs were estimated from Cursor’s token counts; no uncertainty intervals are supplied.
GDPval-AA v2.1: work quality
Artificial Analysis, Stirrup, 220 tasks. Rounded Elo, not percent correct; Index values may differ from later live leaderboards. Max effort; Claude 5.5 and Fable use Default Fallback. Anthropic reports a since-fixed structured-output bug in the prerelease Sonnet deployment used here; these are not confirmed post-fix reruns.
AA-Briefcase v1.1: deliverables
Artificial Analysis: 91 professional-work tasks using Stirrup. Rounded Elo combines rubric, analysis and presentation quality, not completion rate. Max effort; Claude 5.5 and Fable use Default Fallback. The since-fixed prerelease Sonnet structured-output bug also affected this run; a fix does not establish a rerun.
Anthropic-reported evaluation
Humanity’s Last Exam: with tools
Anthropic’s 2,500-question evaluation, five trials at max effort, with web search, fetch and code execution. This is the tool-assisted version, not Artificial Analysis’s separately configured HLE.
Anthropic-reported evaluation
Chartography: with tools
Anthropic: 100 visual-reasoning tasks, five runs at max effort, image cropping and code tools, Gemini 3.5 Flash grader. These are point estimates; the close Sonnet/Opus scores do not establish a decisive lead.
Anthropic-reported evaluation
OSWorld 2.1: partial credit
Anthropic: 108 computer-use tasks, five runs, max effort, 500-action limit and September 10 task/assets versions. Partial credit weights completed checkpoints. An 80.1% score does not mean 80.1% of tasks were fully completed.
Anthropic-reported evaluation
OSWorld 2.1: strict completion
Same Anthropic OSWorld 2.1 setup as the partial-credit chart. Strict pass requires every checkpoint to be satisfied. Keeping the two scores separate shows the gap between partial progress and a fully completed task.
AutomationBench: Zapier run
Zapier v1.0.6, reported by Anthropic. Max-effort task pass rates require all assertions to pass. Sonnet 5.5 used default fallbacks; Opus’s 42.5% includes selective fallback-enabled reruns of refusal tasks (40.0% before reruns). Other rows’ fallback settings are not established here. Not a uniform rerun or AA’s objective-completion metric.
AutomationBench-AA: objectives
DeepSeek V4.1 Flash (max)
68.9%
Artificial Analysis, dataset v1.0.6: 657 private tasks, 50-turn cap. This measures the mean share of guardrail-adjusted objectives completed, not fully completed tasks. Max effort; Sonnet and Opus use Default Fallback.
Anthropic-reported evaluation
Toolathlon-Verified: repeatable success
Anthropic’s internal harness, June 2026 final release: 108 tasks, three trials, adaptive max effort. Pass³ requires all three trials to succeed. Sonnet trails Opus and Fable. Classifiers were on for those three models; two blocked Sonnet trials count as failures. Not the public leaderboard setup.
Choose the effort that fits
Sonnet 5.5: effort vs. Index score
Artificial Analysis Intelligence Index v4.3.2, September 28. Rounded index points across five adaptive-reasoning settings, all with Default Fallback. Max raises the score, but the cost chart shows why it is not an automatic everyday default.
Task cost, not token price
Sonnet 5.5: cost per Index task
Artificial Analysis’s weighted mean Index task-cost estimate, based on token usage, list prices and representative cache-hit rates. Same effort configurations as the score chart. These are evaluation costs, not a price for every real-world task.
Task cost, not token price
Max effort: cost per Index task
Artificial Analysis, all at max effort; Claude 5.5 and Fable use Default Fallback. Sonnet’s weighted mean is $7.60 per Index task versus Opus’s $5.98 despite lower token rates. These averages include unsuccessful attempts, not costs per successful completion. Lower is better.
Published API list prices
API input: per million tokens
Anthropic standard USD rates, September 28, 2026. Sonnet 5.5 keeps Sonnet 5’s token prices; task savings come from usage efficiency. Batch discounts, caching, service modifiers and tool charges are separate. Not SPIRITT subscription prices.
Published API list prices
API output: per million tokens
Anthropic standard USD rates, September 28, 2026. Sonnet 5.5 keeps Sonnet 5’s token prices; task savings come from usage efficiency. Batch discounts, caching, service modifiers and tool charges are separate. Not SPIRITT subscription prices.
Sources: Anthropic’s September 28 launch, system card and API documentation; Artificial Analysis model data and Intelligence Index v4.3.2, checked September 28, 2026. Benchmark versions, reasoning settings, fallbacks and scoring differ. Percentages are rounded to one decimal; Elo and Index values are rounded to integers; cost estimates to cents. Close scores do not establish statistically meaningful superiority. AA marks CritPt, 10% of this Index, under review; these are its published scores, not a recalculated index. Effort names do not equalize compute across models. API list prices and evaluation task costs are not SPIRITT plan prices or universal savings guarantees.