SPIRITT logoSPIRITTAnthropicClaude Opus 5.5

Opus 5.5 does more of the work

Bigger code changes. Deeper research. Clearer answers. Claude Opus 5.5 combines frontier capability with a more efficient default: Anthropic reports 40% lower running costs than Opus 5 on typical workloads. Explore what it can do, then bring your next project to a SPIRITT Workspace.

Opus 5.5 in action

See how people are using Opus 5.5 for apps, simulations, videos, code upgrades, and real business work.

Bigger projects. Fewer detours.

Opus 5.5 pairs stronger coding and knowledge work with fewer detours, clearer communication, and lower standard token prices. Medium is the default effort; higher settings are there when the task needs them.

More progress, less back-and-forth

Anthropic’s launch results put Opus 5.5 ahead on coding, professional knowledge work, chart understanding, and computer use. Early testers report shorter agent sessions, fewer revisions, and clearer explanations on large projects.

The model supports a 1M-token context window and up to 128K output tokens, with text and image input. Standard API rates are $4/M input, $20/M output, and $0.20/M cache reads across the full context window. Input and output rates are 20% lower than Opus 5; cache reads are 60% lower.

The reported 40% saving is a typical-workload result at default settings, not a blanket token discount. Anthropic also reports more than 30% faster output generation. For demanding tasks, compare higher effort against the extra time and cost.

Agentic codingResearch + business work1M contextClearer communicationMedium default effortLower cache-read costs

AA Intelligence Index v4.3.2

Winner
Opus 5.5 max
57.6
Fable 5.1 max
53.4
GPT-6 Astra max
52.7
Opus 5 max
50.8
GPT-5.6 Sol max
47

Artificial Analysis, September 22. All models use max effort. Opus 5.5 and Fable 5.1 use the published default safeguard fallback configuration.

Where Opus 5.5 moves the needle

Compare coding, research, computer use, and cost. The strongest gains sit alongside practical trade-offs in effort, task difficulty, and how success is scored.

Agentic coding
Winner

Terminal-Bench 4.0

Opus 5.5 (xhigh)
66.4%
GPT-6 Astra (high)
57.9%
Fable 5.1 (max)
55.8%
Opus 5 (max)
52.3%
GPT-5.6 Sol (max)
37.3%

Anthropic, September 22. Opus 5.5 uses xhigh; Astra uses high. Opus 5.5 standard error is ±2.6 points. Effort levels are shown in the labels.

Agentic coding
Winner

FrontierCode v1.1 (Main)

Opus 5.5 (max)
54.4%
GPT-6 Astra (max)
53.3%
Fable 5.1 (max)
50.3%
Opus 5 (max)
48%
GPT-5.6 Sol (max)
47.5%

Anthropic launch comparison at max effort. The default medium setting scored 54.6% in the separate effort curve, showing why more thinking is not always better.

Agentic coding
Winner

CursorBench 4.0

Opus 5.5 (max)
57.8%
Fable 5.1 (max)
51.8%
Opus 5 (max)
46.6%
GPT-5.6 Sol (max)
41.7%

Ambiguous, multi-file coding tasks in Cursor. Max effort for the displayed models. The default Opus 5.5 medium setting scores 52.5%.

Knowledge work
Winner

GDPval-AA v2.1

Opus 5.5 (max)
1846
Fable 5.1 (max)
1735
Opus 5 (max)
1708
GPT-5.6 Sol (max)
1588
GPT-6 Astra (max)
1542

Professional work across 44 occupations, scored in Elo. These are the max-effort values in Anthropic’s launch comparison.

Business workflows
Near frontier

AutomationBench

GPT-6 Astra (max)
41.4%
Opus 5.5 (max)
40%
Fable 5.1 (max)
31.4%
GPT-5.6 Sol (max)
28.8%
Opus 5 (max)
26.9%

Zapier’s business-workflow evaluation. All displayed models use max effort. Refusals count as failures; Astra leads this comparison.

Multidisciplinary reasoning
Winner

Humanity's Last Exam (with tools)

Opus 5.5 (max)
67.7%
Fable 5.1
65.6%
Opus 5
63.6%
GPT-6 Astra
57.2%

Anthropic’s multimodal evaluation with web tools and code execution. Opus 5.5 uses max effort; other efforts are not specified in this table.

Agentic scientific research
Near frontier

Terminal-Bench-Science 0.1

GPT-6 Astra (max)
64.6%
Opus 5.5 (max)
58.7%
Fable 5.1 (max)
52.6%
Opus 5 (max)
29%
GPT-5.6 Sol
22.4%

Scientific research in a terminal environment. Astra leads. Reported uncertainty is approximately ±3.5–5 points per model.

Computer use
Winner

OSWorld 2.0 (partial)

Opus 5.5 (max)
81.8%
Fable 5.1 (max)
80.7%
Opus 5 (max)
74%

Partial-credit computer-use scores at max effort. Opus 5.5 receives credit for completed checkpoints; full-task success is shown separately below.

Visual chart recognition
Winner

Chartography (with tools)

Opus 5.5 (max)
89%
Fable 5.1 (max)
88.4%
Opus 5 (max)
83.4%

Chart-reading performance with tools at max effort, from Anthropic’s launch report. Tool-assisted scores should be compared with the same setup.

Research
Winner

WANDR: large-scale research

Opus 5.5 (max)
72.3%
Fable 5.1 (max)
68.7%
Opus 5 (max)
67.2%

Large-scale data collection, scored with soft F1. Anthropic uses offline web tools and a 980K task-token budget; compare these rows with one another.

Computer use
Winner

OSWorld 2.0: full-task success

Opus 5.5 max
48.7%
Fable 5.1 max
42.8%
Opus 5 max
37.2%

All task checkpoints must pass. These stricter success rates are different from the partial-credit scores above. Source: Opus 5.5 System Card.

Visual understanding
Winner

Chartography without tools

Opus 5.5 max
64.4%
Fable 5.1 max
44.8%
Opus 5 max
29.8%

Chart recognition without tool assistance. This highlights the model’s direct visual reasoning rather than the higher tool-assisted scores. Source: Anthropic System Card.

Coding economics
Winner

FrontierCode: cost at strong settings

Opus 5.5 medium
$0.80
Fable 5.1 low
$2.47
GPT-6 Astra max
$4.36
Opus 5 medium
$4.61

These selected settings score 52.8–54.6% on FrontierCode Main. Effort levels differ as shown; costs are Anthropic’s estimates per task.

Published API pricing
Winner

Input price per million tokens

Opus 5.5
$4.00
Opus 5
$5.00

Anthropic’s standard global API rates. Input tokens cost 20% less than Opus 5. These are token prices, not a SPIRITT subscription price.

Published API pricing
Winner

Output price per million tokens

Opus 5.5
$20.00
Opus 5
$25.00

Anthropic’s standard global API rates. Output tokens cost 20% less; total workload savings also depend on token usage and effort.

Agent-workflow economics
Winner

Cache-read price per million tokens

Opus 5.5
$0.20
Opus 5
$0.50

Repeated context costs 60% less to read from cache than with Opus 5. Caching is especially relevant to long coding and agent sessions.

Independent cost trade-off

AA cost per task at max effort

GPT-5.6 Sol max
$1.99
GPT-6 Astra max
$3.26
Opus 5 max
$5.86
Opus 5.5 max
$5.98
Fable 5.1 max
$7.63

Max effort reaches higher capability but is not the cheapest way to run every task. In AA’s same-index snapshot, Opus 5.5 medium costs $1.34 per task.

Sources: Anthropic’s September 22 release and System Card, Artificial Analysis v4.3.2, and CursorBench. Anthropic’s Opus 5.5 results generally use max adaptive effort, with xhigh on Terminal-Bench 4.0. Safeguard-triggered fallback is included where enabled; Zapier AutomationBench treats refusals as failures. Partial-credit and full-task success rates are labeled separately. API prices are Anthropic’s list prices; tools, special service tiers, and regional options can add charges.

How It Works

Give the project a workspace, choose the model setup, and work toward a concrete finish line.

01

Bring the project into SPIRITT

Start with your repository, documents, research question, or business goal. Your workspace keeps the browser, files, terminal, integrations, and memory together so your agent can do real work.

A SPIRITT workspace with browser, files, terminal, and connected tools
02

Choose Claude for the task

Check the Claude models and connection options available in your workspace. Opus 5.5’s medium default is designed for everyday work; raise the effort when a harder problem justifies it.

The Claude sunburst in translucent orange glass above a pastel plinth
03

Build, check, and ship

Ask for the finished outcome, not just an answer: a tested code change, a usable app, or a report with supporting sources. Keep the outputs and decisions together, and review important actions before they go live.

Connected tools for building, checking, and shipping work

Turn the next big ask into finished work

Put your agent in a workspace with the tools, files, and memory to carry a project through.

Questions

Frequently asked questions

01What is new in Claude Opus 5.5?+
Opus 5.5 is the first release in Anthropic’s Claude 5.5 family. It improves agentic coding, professional knowledge work, computer use, and communication while lowering standard API token rates. It is designed to stay with larger tasks and explain its work more clearly.
02Is it really 40% cheaper than Opus 5?+
Anthropic reports 40% lower costs on typical workloads at default settings, combining lower rates with more efficient task completion. Input and output tokens are 20% cheaper, and cache reads are 60% cheaper. A particular task can cost more or less, especially at higher effort.
03What are the API prices?+
Per million tokens, standard global prices are $4 input, $20 output, and $0.20 cache reads. Five-minute cache writes cost $5 and one-hour writes cost $8. Batch processing discounts input and output by 50%. Standard rates apply across the 1M context window; tools and optional service settings can add charges. These are Anthropic API prices, not SPIRITT plan prices.
04Should I use max effort for everything?+
No. Medium is the default and is a strong starting point. Higher effort can improve difficult work, but uses more time and tokens. Artificial Analysis’s launch snapshot puts Opus 5.5 medium around $1.34 per task and max around $5.98, at different capability levels. Test the setting that fits your actual work.
05What is Fast mode?+
Fast mode uses the same model with higher output throughput, advertised at up to 2.5 times standard speed. It costs $8/M input and $40/M output. Claude API access is a research preview requiring approval, and Fast mode is not compatible with Batch API. Higher output speed does not guarantee faster first-token or total task time.
06Does Opus 5.5 win every comparison?+
No. In the launch table, GPT-6 Astra leads AutomationBench and Terminal-Bench-Science. Some System Card tasks also favor earlier models. Benchmark versions, tool access, effort and scoring rules matter, so use the comparisons as a starting point for your own workload.
07How do the safeguards affect work?+
Anthropic reports stronger prompt-injection resistance and better respect for action boundaries. In supported products and opted-in API configurations, sensitive cybersecurity requests can fall back to Opus 4.8, while some biology and frontier-model-development requests can fall back to Opus 5. Important conclusions and irreversible actions still need appropriate review.
08Can it work with images, video, and large documents?+
Its native inputs are text and images, and its native output is text. It supports vision, PDFs, tool use, and a 1M-token context window. The creative projects shown above use the model with software and tools; native audio or video generation is not listed as a model capability.
09How do I check Opus 5.5 access in my workspace?+
Open your SPIRITT Workspace and check its enabled models and connected Claude options. Model availability can depend on the workspace configuration and access route. Choose an available setup before starting a model-specific workflow.
Buy from builders who use what they sellBuilt usingSPIRITT