SPIRITT logoSPIRITTAnthropicClaude Sonnet 5.5

Sonnet 5.5. Keep the work moving.

Fix the bug. Refine the design. Finish the deck. Claude Sonnet 5.5 focuses on well-scoped everyday work, with Anthropic reporting 30%+ faster output generation and up to 30% lower task costs than Sonnet 5.

Sonnet 5.5 in action

Explore interactive builds, coding demos, and early evaluations of Sonnet 5.5, alongside Anthropic’s official launch.

More progress on the work you do every day.

Sonnet 5.5 is built for focused coding, fast iteration, and polished documents, slides, and spreadsheets. Keep Opus 5.5 for the most complex work that demands sustained judgment.

A faster everyday complement to Opus

Anthropic introduced Claude Sonnet 5.5 on September 28, 2026. It supports text and image input, text output, a one-million-token context window, and up to 128,000 output tokens in standard requests. Adaptive reasoning offers five effort levels, from low to max.

Standard API rates are $2 per million input tokens, $10 per million output tokens, and $0.20 per million cache-read tokens. These are unchanged from Sonnet 5. Anthropic’s claimed savings come from using fewer tokens on typical work, not a token-price cut.

At max effort, Sonnet 5.5 scores 56 on Artificial Analysis’s Intelligence Index v4.3.2, close to Opus 5.5’s 58. But max is not always the economical setting: AA estimates $7.60 per Index task for Sonnet versus $5.98 for Opus at max. Match the effort to the job.

Anthropic reports major coding gains and stronger visual understanding, while noting that Opus 5.5 remains better at complex, open-ended work. Model access in SPIRITT depends on the models and connections enabled for your workspace.

1M context128K standard outputText + image inputFive effort levels$2 / $10 API pricingRelease brief

Artificial Analysis Intelligence Index

Opus 5.5 (max)
58
Opus 5.5 (xhigh)
56
Sonnet 5.5 (max)
56
Fable 5.1 (max)
53
GPT-6 Astra (max)
53
Sonnet 5 (max)
38

Index v4.3.2, September 28, 2026. Selected configurations; rounded index points, not percentages. Claude uses adaptive reasoning; Sonnet 5.5, Opus 5.5 and Fable 5.1 use Default Fallback. Sonnet max and Opus xhigh both round to 56; their tiny unrounded gap does not establish a meaningful difference.

Compare the results. Choose the right effort.

Selected coding, knowledge-work, visual-reasoning, and automation results, plus the cost of reaching them. Independent Artificial Analysis runs stay separate from Anthropic’s reported evaluations.

Anthropic-reported evaluation

Terminal-Bench 4.0: Anthropic

Sonnet 5.5 (max)
70.6%
Opus 5.5 (xhigh)
66.4%
Fable 5.1 (max)
55.8%
Sonnet 5 (max)
10.3%

Anthropic’s Claude Code bare-mode run: 66 tasks, cached dependencies, no internet egress. Sonnet and Opus 5.5 used five trials per task with safeguards; fallback affected 1.2% and 2.5% of requests respectively. Sonnet is max; Opus is xhigh. Other comparators’ repeat counts differ. Separate from AA’s run.

Artificial Analysis

Terminal-Bench 4.0: AA

Sonnet 5.5 (max)
63.6%
Opus 5.5 (max)
59.6%
GPT-6 Astra (max)
59.1%
Fable 5.1 (max)
52%
Sonnet 5 (max)
14.1%

Artificial Analysis: mini-swe-agent, 66 tasks, three repeats, pass@1, 500-step cap. All rows use max effort; Claude 5.5 configurations and Fable use Default Fallback. These results must not be substituted for Anthropic’s Claude Code run.

Cognition, via Anthropic

FrontierCode v1.1 Main

Opus 5.5 (max)
54.4%
Sonnet 5.5 (xhigh)
52.1%
GPT-6 Sol (max)
49.3%
Sonnet 5.5 (max)
46.2%
Sonnet 5 (max)
42.4%

Cognition’s 150-task evaluation, reported by Anthropic, using Claude Code and Codex CLI. Composite scores, not simple test pass rates. Sonnet scores lower at max than xhigh. In two cases Cognition examined, extra code-review agents caused a timeout or out-of-scope edits.

Cursor, via Anthropic

CursorBench 4.0

Opus 5.5 (max)
57.8%
Sonnet 5.5 (max)
55.5%
Sonnet 5 (max)
34.1%

Cursor’s production-agent evaluation, reported by Anthropic. Max-effort scores shown. Sonnet’s reported task costs were estimated from Cursor’s token counts; no uncertainty intervals are supplied.

Artificial Analysis

GDPval-AA v2.1: work quality

Opus 5.5 (max)
1846
Sonnet 5.5 (max)
1844
Fable 5.1 (max)
1735
GPT-6 Astra (max)
1542
Sonnet 5 (max)
1449

Artificial Analysis, Stirrup, 220 tasks. Rounded Elo, not percent correct; Index values may differ from later live leaderboards. Max effort; Claude 5.5 and Fable use Default Fallback. Anthropic reports a since-fixed structured-output bug in the prerelease Sonnet deployment used here; these are not confirmed post-fix reruns.

Artificial Analysis

AA-Briefcase v1.1: deliverables

Opus 5.5 (max)
1822
Sonnet 5.5 (max)
1811
Fable 5.1 (max)
1678
Sonnet 5 (max)
1359

Artificial Analysis: 91 professional-work tasks using Stirrup. Rounded Elo combines rubric, analysis and presentation quality, not completion rate. Max effort; Claude 5.5 and Fable use Default Fallback. The since-fixed prerelease Sonnet structured-output bug also affected this run; a fix does not establish a rerun.

Anthropic-reported evaluation

Humanity’s Last Exam: with tools

Opus 5.5 (max)
67.7%
Sonnet 5.5 (max)
64.5%
Sonnet 5 (max)
54.9%

Anthropic’s 2,500-question evaluation, five trials at max effort, with web search, fetch and code execution. This is the tool-assisted version, not Artificial Analysis’s separately configured HLE.

Anthropic-reported evaluation

Chartography: with tools

Sonnet 5.5 (max)
90.2%
Opus 5.5 (max)
89%
Fable 5.1 (max)
88.4%

Anthropic: 100 visual-reasoning tasks, five runs at max effort, image cropping and code tools, Gemini 3.5 Flash grader. These are point estimates; the close Sonnet/Opus scores do not establish a decisive lead.

Anthropic-reported evaluation

OSWorld 2.1: partial credit

Opus 5.5 (max)
81.8%
Sonnet 5.5 (max)
80.1%
Sonnet 5 (max)
57%

Anthropic: 108 computer-use tasks, five runs, max effort, 500-action limit and September 10 task/assets versions. Partial credit weights completed checkpoints. An 80.1% score does not mean 80.1% of tasks were fully completed.

Anthropic-reported evaluation

OSWorld 2.1: strict completion

Opus 5.5 (max)
48.7%
Sonnet 5.5 (max)
43.5%
Sonnet 5 (max)
25.6%

Same Anthropic OSWorld 2.1 setup as the partial-credit chart. Strict pass requires every checkpoint to be satisfied. Keeping the two scores separate shows the gap between partial progress and a fully completed task.

Zapier, via Anthropic

AutomationBench: Zapier run

Sonnet 5.5 (max)
44.7%
Opus 5.5 (max)
42.5%
GPT-6 Sol (max)
32%
Sonnet 5 (max)
10.7%

Zapier v1.0.6, reported by Anthropic. Max-effort task pass rates require all assertions to pass. Sonnet 5.5 used default fallbacks; Opus’s 42.5% includes selective fallback-enabled reruns of refusal tasks (40.0% before reruns). Other rows’ fallback settings are not established here. Not a uniform rerun or AA’s objective-completion metric.

Artificial Analysis

AutomationBench-AA: objectives

Sonnet 5.5 (max)
71.3%
Opus 5.5 (max)
69.5%
DeepSeek V4.1 Flash (max)
68.9%
GPT-6 Astra (max)
68.5%

Artificial Analysis, dataset v1.0.6: 657 private tasks, 50-turn cap. This measures the mean share of guardrail-adjusted objectives completed, not fully completed tasks. Max effort; Sonnet and Opus use Default Fallback.

Anthropic-reported evaluation

Toolathlon-Verified: repeatable success

Fable 5.1
73.1%
Opus 5.5
72.2%
Sonnet 5.5
68.5%
Sonnet 5
65.7%

Anthropic’s internal harness, June 2026 final release: 108 tasks, three trials, adaptive max effort. Pass³ requires all three trials to succeed. Sonnet trails Opus and Fable. Classifiers were on for those three models; two blocked Sonnet trials count as failures. Not the public leaderboard setup.

Choose the effort that fits

Sonnet 5.5: effort vs. Index score

Sonnet 5.5 (max)
56
Sonnet 5.5 (xhigh)
52
Sonnet 5.5 (high)
47
Sonnet 5.5 (medium)
41
Sonnet 5.5 (low)
36

Artificial Analysis Intelligence Index v4.3.2, September 28. Rounded index points across five adaptive-reasoning settings, all with Default Fallback. Max raises the score, but the cost chart shows why it is not an automatic everyday default.

Task cost, not token price

Sonnet 5.5: cost per Index task

Sonnet 5.5 (low)
$0.41
Sonnet 5.5 (medium)
$0.59
Sonnet 5.5 (high)
$1.08
Sonnet 5.5 (xhigh)
$2.74
Sonnet 5.5 (max)
$7.60

Artificial Analysis’s weighted mean Index task-cost estimate, based on token usage, list prices and representative cache-hit rates. Same effort configurations as the score chart. These are evaluation costs, not a price for every real-world task.

Task cost, not token price

Max effort: cost per Index task

GPT-6 Astra (max)
$3.26
Sonnet 5 (max)
$5.09
Opus 5.5 (max)
$5.98
Sonnet 5.5 (max)
$7.60
Fable 5.1 (max)
$7.63

Artificial Analysis, all at max effort; Claude 5.5 and Fable use Default Fallback. Sonnet’s weighted mean is $7.60 per Index task versus Opus’s $5.98 despite lower token rates. These averages include unsuccessful attempts, not costs per successful completion. Lower is better.

Published API list prices

API input: per million tokens

Sonnet 5.5
$2.00
Sonnet 5
$2.00
Opus 5.5
$4.00

Anthropic standard USD rates, September 28, 2026. Sonnet 5.5 keeps Sonnet 5’s token prices; task savings come from usage efficiency. Batch discounts, caching, service modifiers and tool charges are separate. Not SPIRITT subscription prices.

Published API list prices

API output: per million tokens

Sonnet 5.5
$10.00
Sonnet 5
$10.00
Opus 5.5
$20.00

Anthropic standard USD rates, September 28, 2026. Sonnet 5.5 keeps Sonnet 5’s token prices; task savings come from usage efficiency. Batch discounts, caching, service modifiers and tool charges are separate. Not SPIRITT subscription prices.

Sources: Anthropic’s September 28 launch, system card and API documentation; Artificial Analysis model data and Intelligence Index v4.3.2, checked September 28, 2026. Benchmark versions, reasoning settings, fallbacks and scoring differ. Percentages are rounded to one decimal; Elo and Index values are rounded to integers; cost estimates to cents. Close scores do not establish statistically meaningful superiority. AA marks CritPt, 10% of this Index, under review; these are its published scores, not a recalculated index. Effort names do not equalize compute across models. API list prices and evaluation task costs are not SPIRITT plan prices or universal savings guarantees.

Turn everyday work into finished work.

Bring a clear goal, use the model access available in your workspace, and judge the finished result.

01

Bring one clear task

Start with a bug to fix, a page to improve, or a document to finish. Give SPIRITT the relevant files and a clear definition of what a good result looks like.

A soft 3D illustration of a workspace for a focused project
02

Check your Sonnet 5.5 access

Check whether Sonnet 5.5 is available through your enabled models or connections. Access varies by workspace. Where it is enabled, start with a focused task and compare the result against your usual model.

A peach-orange glass Claude-inspired asterisk floating on an ivory background
03

Tune effort to the job

Try a representative task before scaling up. Compare quality, total cost and completion time, then raise the effort where the result benefits. For complex open-ended judgment, compare against Opus 5.5.

A soft 3D illustration of an idea becoming a finished project

Keep the next project moving.

Bring your goal into SPIRITT and work with the models available in your workspace. Try Sonnet 5.5 on the tasks that matter to you wherever access is enabled.

Questions

A clearer picture before you start.

01What is Claude Sonnet 5.5?+
Claude Sonnet 5.5 is Anthropic’s September 28, 2026 model for well-scoped everyday tasks, coding, design iteration, documents, slides and spreadsheets. Anthropic positions it as a faster, lower-cost complement to Opus 5.5, not a replacement for Opus on complex, open-ended work.
02What does Sonnet 5.5 cost?+
Anthropic’s standard API list prices are $2 per million input tokens and $10 per million output tokens. Cache reads cost $0.20 per million tokens; five-minute cache writes cost $2.50 and one-hour writes $4. Batch input/output rates are $1/$5. Standard rates cover the full one-million-token context. Service and residency modifiers or tool charges may apply. These are API prices, not SPIRITT subscription prices.
03Is Sonnet 5.5 really 30% cheaper and faster?+
Anthropic reports 30%+ faster output generation and up to 30% lower cost per task than Sonnet 5 in its testing. Per-token prices have not changed. Fewer tokens can make a task cheaper, but workload and reasoning effort matter. Artificial Analysis’s max-effort Index task cost is $7.60 for Sonnet 5.5 versus $5.98 for Opus 5.5, so savings are not universal. Output generation speed is not a guarantee of total task time.
04Should I choose Sonnet 5.5 or Opus 5.5?+
Start by comparing Sonnet on focused tasks and quick iterations, especially when responsiveness and token rates matter. Anthropic still recommends Opus for complex, open-ended work requiring sustained judgment. Similar scores on a particular benchmark do not make the two models interchangeable, and lower token rates do not always mean lower task costs.
05Which reasoning effort should I use?+
Sonnet 5.5 offers low, medium, high, xhigh and max effort. Anthropic’s documented defaults are high for the API and medium in Claude apps and Claude Code. Levels were recalibrated from Sonnet 5, so the same name does not imply the same thinking budget. Match effort to the task rather than always choosing max: in the reported FrontierCode run, Sonnet scored 52.1% at xhigh but 46.2% at max.
06What are the context and output limits?+
Standard requests support a one-million-token context window and up to 128,000 output tokens. Anthropic separately documents a 300,000-token output beta for batch requests; it is not the standard output limit and requires the relevant beta configuration.
07Can Sonnet 5.5 work with images, audio or video?+
The documented model supports text and image inputs and text output, including code. Image understanding is different from native image, audio or video generation. A video showing an app or animation built with the model does not mean the model directly generates video files.
08Why are there two Terminal-Bench 4.0 results?+
Anthropic reports 70.6% for its Claude Code setup. Artificial Analysis reports about 63.6% using mini-swe-agent. They use different evaluation setups and repeat counts. Neither number should replace the other, and both are distinct from older Terminal-Bench versions.
09Does the 80.1% OSWorld score mean full task completion?+
No. Anthropic’s 80.1% is partial credit across checkpoints. Its strict-completion score is 43.5%, requiring every checkpoint to pass. The page shows both because partial progress and finishing the entire task are different outcomes.
10Why do the automation scores differ?+
The Zapier run and AutomationBench-AA use different metrics. AA reports the average fraction of guardrail-adjusted objectives completed, not the share of tasks fully finished. Opus’s 42.5% in the reported Zapier run also includes fallback-enabled reruns of refusal tasks; it is not the earlier 40.0% under the original conditions.
11What limitations should I keep in mind?+
More reasoning is not always better or cheaper, and the system card documents regressions on some safety and reliability tests. Cyber safeguards can refuse or fall back on higher-risk requests; automatic fallback is not universal across API integrations. Review outputs and test representative workflows. Existing API integrations may need changes to thinking, tool use or sampling settings.
12Where do these benchmarks come from?+
The page combines clearly labeled Anthropic-reported results, including evaluations by Cognition, Cursor and Zapier, with separately sourced Artificial Analysis results. AA Index-contribution Elo can remain frozen while live leaderboards change. Anthropic reports a since-fixed structured-output bug in the prerelease Sonnet 5.5 deployment used for GDPval-AA and AA-Briefcase. These recorded scores are not confirmed post-fix reruns. Small point-estimate gaps are not proof of a decisive advantage.
13Can I use Claude Sonnet 5.5 in SPIRITT?+
Access depends on the models and connections enabled for your workspace. This page covers Sonnet 5.5’s release and published results; it does not promise the model is already enabled in every SPIRITT workspace. Open a workspace to start with the available models and check your Sonnet 5.5 access options.
Buy from builders who use what they sellBuilt usingSPIRITT