Terminal-Bench 4.0
Anthropic, September 22. Opus 5.5 uses xhigh; Astra uses high. Opus 5.5 standard error is ±2.6 points. Effort levels are shown in the labels.
Bigger code changes. Deeper research. Clearer answers. Claude Opus 5.5 combines frontier capability with a more efficient default: Anthropic reports 40% lower running costs than Opus 5 on typical workloads. Explore what it can do, then bring your next project to a SPIRITT Workspace.
See how people are using Opus 5.5 for apps, simulations, videos, code upgrades, and real business work.
Claude Opus 5.5 is available today.
— Anthropic (@AnthropicAI) 2026-09-22T16:31:47.000Z
Introducing Claude Opus 5.5, the first model in our new Claude 5.5 family. It performs at the level of Claude Fable 5.1 for most tasks, and costs 40% less to run than Opus 5. https://t.co/Q9C2VKQ79f
— Claude (@claudeai) 2026-09-22T16:31:01.000Z
sketch-to-simulation with Opus 5.5 https://t.co/XPdM1kCRTe
— Ben Poole (@poolio) 2026-09-22T17:11:17.000Z
Claude Opus 5.5 1-shotted the launch video of @shotbaseapp using @HyperFrames_ in less than 20 mins no way 🤯🤯🤯🤯 https://t.co/U20jJkHn41
— Miguel Ángel (@Miguel07Code) 2026-09-22T16:55:39.000Z
Ran the exact same prompt in Opus 5.5! Let’s see how @AnthropicAI’s model holds up here 🔍👇 https://t.co/cwGTNpPnas
— Verdent (@verdent_ai) 2026-09-22T17:12:52.000Z
Anthropic gave me early access to Opus 5.5 (no idea) I will say not knowing what it is changes your opinion on it: is it Fable-like? Opus-like? I actually think playing a little model lottery might help you shape your opinions a bit better. I ran a couple large tasks with it just to see how well it'd perform, and it did well enough that I could barely tell the difference between it and Fable. The first was a set of data analysis/curation for Peated. The second large task was a big migration in Junior. The behavior was different enough that it was noticeable. It reached for subagents more, but also made less assumptions (this is a pro and con). It performed well enough that I had to go dig into a comparable of it vs Fable to even determine which one had better outputs (both, neither). They both created quality outcomes, but the tradeoffs they made were different in each. My feedback was pretty clear: given how well its performing, if you told me this was cheaper than Fable currently is I'd be really happy. Performed a lot better than Opus 5 in my testing and given today's announcement I now know it is cheaper than both Opus 5 and Fable. Seems like a great improvement, and has gotten me to give Anthropic models another shot. Timing is probably good as I just ran out of my Fable budget this morning...
— David Cramer (@zeeg) 2026-09-22T16:42:04.000Z
Opus 5.5 is a really good model. It's been my daily driver the last few weeks. We had Opus 5.5 and Fable 5.1 each port HAProxy from C to Rust. Both passed nearly all of HAProxy's tests, but Opus 5.5 finished in 9.5 hours compared to Fable 5.1's 12 hours, and for 51% less cost.
— Boris Cherny (@bcherny) 2026-09-22T16:45:10.000Z
Claude Opus 5.5 is on Google Cloud! Built for everyday complex tasks, it handles long-running coding and knowledge work while reporting back clearly on what it did, what it found, and what it needs next at a lower cost per task. Try it in Model Garden → https://t.co/seLCbcFygF https://t.co/8hU2YplZVb
— Google Cloud Tech (@GoogleCloudTech) 2026-09-22T16:58:22.000Z
We evaluated @claudeai Opus 5.5 on Box's Complex Work Eval. The standout finding is efficiency. → 30% faster than Opus 5 → A third of the tokens → Answers 42% shorter with the findings intact What disappeared was restatement and preamble. Opus 5.5 simply led with the conclusions. Accuracy held across Consumer Products, Financial Services, and Technology. Claude Opus 5.5 is coming to Box AI soon.
— Box (@Box) 2026-09-22T16:42:31.000Z
Claude Opus 5.5 is here! We ran the numbers. Here's where you should upgrade from Opus 5 👇️ Any crucial Ops workflows - it scored 62% for Ops at Max effort, at roughly $1.37 per task. Great value for the price. Overall it scored 40% across 657 tasks - a BIG jump from Opus 5's 26.9%, and one of the best models we've benchmarked yet. Note: We'd recommend scaling the effort per workflow (instead of keeping the default, Medium, for each). For example, Low scored 48% for Ops compared to Max’s 62%, but Max is double the price. Worth experimenting! See all the models we've benchmarked here: https://t.co/CHiZZRCTuR
— Zapier (@zapier) 2026-09-22T16:42:45.000Z
Opus 5.5 is live in Factory. Some initial observations: / Medium is a strong default / 20–25% fewer output tokens than @AnthropicAI Opus 5 at the same effort / Clear, actionable answers on long investigations Try it now: https://t.co/fjCFt0ZfyV https://t.co/pZN8krlWEa
— Factory (@FactoryAI) 2026-09-22T16:48:27.000Z
Claude Opus 5.5 is now available to all Perplexity Computer users. On our Wide-And-Deep-Research evals, it compares favorably to Fable 5.1 at a fraction of the cost. Opus 5.5 will become the “Standard” Effort orchestrator on Computer for all Pro and Max users. Congrats @AnthropicAI!
— Aravind Srinivas (@AravSrinivas) 2026-09-22T16:49:19.000Z
Claude Opus 5.5 takes the top spot on the Artificial Analysis Intelligence Index, along with a 20% price cut and larger cache hit discount Claude Opus 5.5 brings Anthropic to parity with GPT-6 Astra on evaluations like Terminal-Bench 4.0 and AutomationBench-AA, while extending Anthropic’s lead in agentic knowledge work. At max effort it scores 58 on the Artificial Analysis Intelligence Index, the highest score we have measured by several points. Anthropic has cut Opus pricing to $4/$20 per 1M input/output tokens (Opus 5: $5/$25) and cache reads from $0.50 to $0.20. Key takeaways: ➤ Consistent strong performance, with leading scores on six of the ten Intelligence Index evaluations: Humanity's Last Exam 61.4% (previous best 59.1%, Claude Fable 5.1), SciCode 66.9% (63.1%, Fable 5.1), GDPval-AA v2.1, AA-Briefcase v1.1, AA-Omniscience and AutomationBench-AA. On Terminal-Bench 4.0 it scores 59.6%, level with the leader GPT-6 Astra (xhigh) and +11 points over Opus 5. It remains slightly behind on CritPt, AA-LCR, and GDP.pdf ➤ Leads in agentic knowledge work: On AA-Briefcase, our private frontier knowledge work evaluation, it reaches an Elo of 1822. This is +143 over Fable 5.1, ahead on both analytical quality and presentation, and is the first time Anthropic has reached presentation quality surpassing GPT-5.6 Sol. This evaluation tests whether models can produce accurate and well-presented professional outputs using our open source reference agent harness, Stirrup ➤ Level with Opus 5 on cost per task despite 1.6x the output tokens: Opus 5.5 (max) uses ~119k output tokens per Intelligence Index task, against ~73k for Opus 5 (max), ~78k for Fable 5.1 (max) and ~27k for GPT-6 Astra (max) ➤ Four of five effort levels sit on the Intelligence vs Cost per Task frontier: Opus 5.5 max, xhigh, high, and medium all sit on the Pareto frontier, costing less or outperforming other models scoring 50+ (GPT-6 Astra, Claude Fable 5.1, and Claude Opus 5) Other model details: ➤ Context window: 1 million token context with image and text input support, unchanged from Opus 5 ➤ Pricing: $4/$20 per 1M input/output tokens, down 20% from $5/$25 for Opus 5. Cache writes $5 per 1M tokens for the 5 minute TTL, down from $6.25. Cache reads have been further discounted to $0.20 per 1M tokens, down 60% from Opus 5’s $0.50. This is a 95% discount compared to uncached input pricing, up from 90% on previous Opus models ➤ Effort settings: Five effort settings (low, medium, high, xhigh, and max). Intelligence Index evaluations were run at all five with Anthropic's default fallback enabled
— Artificial Analysis (@ArtificialAnlys) 2026-09-22T16:41:45.000Z
Claude Opus 5.5 takes #1 on RSI Index and is the first model to beat the published reference on LM Training under our protocol, marking a major step forward for long-horizon agentic work. https://t.co/gBRuqe4Gzu
— Vals AI (@ValsAI) 2026-09-22T17:10:25.000Z
We ran Opus 5.5 through CodeRabbit's review pipeline. > On 80 known bug patterns it caught 51 vs 49 for our production mix. > On 13 harder cases, 10 vs 5. > It found a retry-count race in Cal.com that production missed. The catch is ~50% more tokens, and 9 bugs our baseline caught that Opus 5.5 didn't. Read our full Opus 5.5 model evaluation here: https://t.co/8GRec6koER
— CodeRabbit (@coderabbitai) 2026-09-22T16:37:01.000Z
BREAKING: Anthropic just dropped Opus 5.5—and it’s pulling some of our recent Codex converts back to Claude. We’ve been testing it at @every across coding, design, writing, and knowledge work. It sometimes beats Fable 5.1 in our testing and is up to 40% cheaper than Opus 5. It's a strong contender for new daily driver model: It’s excellent at end-to-end builds that match your taste. I’ve started reaching for Opus over Fable on big, end-to-end coding projects. And @kieranklaassen is replacing Fable 5.1 with Opus for his day-to-day product work. He calls it his new favorite model. It produces legible prose, but still trails Astra on writing tasks. Opus 5.5 scores a 68.42 on reading ease—the highest on any model we've tested. But in my writing benchmark tests it consistently buries the main point in intro paragraphs, and revisions. You'll be able to understand what this model is saying (yay!) but for day-to-day writing it's still behind. The economics are striking. Anthropic says Opus 5.5 will cost $5 per million input tokens and $20 per million output tokens. That’s the same input price and 20% less for output than Opus 5’s $5/$25. Altogether it should save roughly 40% on costs than Opus 5. It still has a “do the most” problem. @hammermt tested it on our standard knowledge work benchmarks, and it's results were great when thye came back. On at least one test it ran past the 10 minute time limit before delivering. Net Result: If you build apps and interfaces, try it. It’s become my go-to for ambitious coding projects. I’m still roughly 80/20 Codex versus Claude in day-to-day use, and I still prefer Sol and Astra for editing. But I’m spending far more of my tokens with Claude than I was a week ago. State of Play: Codex is still the better harness for me, but Anthropic is steadily gaining ground. They have a history of making their smaller models perform better than their bigger ones (Sonnet 3.7 for example) and they seem to have done the same with Opus 5.5. Full vibe check will be on @every soon!
— Dan Shipper (@danshipper) 2026-09-22T16:32:27.000Z
Opus 5.5 was a good model in my early tests, first non-Fable/Astra model to feel like a Fable-class model, but still hasn't fully solved the dense language issue of the recent Claudes. Its version of the same shader (broken towers are a nice touch): https://t.co/GH5cwxCwu8 https://t.co/KDDQolLSfP
— Ethan Mollick (@emollick) 2026-09-22T16:55:20.000Z
Opus 5.5 on the Bach Benchmark. I’m surprised. There are spacing issues, awkward doublings, and one case of parallel octaves. I think Grok 4.7’s chorale from yesterday might have been better. Prompt: “In LilyPond (version 2.24), write a 4-part chorale in the style of Bach, 3/4 time, G minor. Use two staves -- soprano and alto on the top staff, tenor and bass on the bottom staff. Respond with only the code block.”
— Auggie (@aug5thmusic) 2026-09-22T17:12:47.000Z
Opus 5.5 pairs stronger coding and knowledge work with fewer detours, clearer communication, and lower standard token prices. Medium is the default effort; higher settings are there when the task needs them.
Anthropic’s launch results put Opus 5.5 ahead on coding, professional knowledge work, chart understanding, and computer use. Early testers report shorter agent sessions, fewer revisions, and clearer explanations on large projects.
The model supports a 1M-token context window and up to 128K output tokens, with text and image input. Standard API rates are $4/M input, $20/M output, and $0.20/M cache reads across the full context window. Input and output rates are 20% lower than Opus 5; cache reads are 60% lower.
The reported 40% saving is a typical-workload result at default settings, not a blanket token discount. Anthropic also reports more than 30% faster output generation. For demanding tasks, compare higher effort against the extra time and cost.
Artificial Analysis, September 22. All models use max effort. Opus 5.5 and Fable 5.1 use the published default safeguard fallback configuration.
Compare coding, research, computer use, and cost. The strongest gains sit alongside practical trade-offs in effort, task difficulty, and how success is scored.
Anthropic, September 22. Opus 5.5 uses xhigh; Astra uses high. Opus 5.5 standard error is ±2.6 points. Effort levels are shown in the labels.
Anthropic launch comparison at max effort. The default medium setting scored 54.6% in the separate effort curve, showing why more thinking is not always better.
Ambiguous, multi-file coding tasks in Cursor. Max effort for the displayed models. The default Opus 5.5 medium setting scores 52.5%.
Professional work across 44 occupations, scored in Elo. These are the max-effort values in Anthropic’s launch comparison.
Zapier’s business-workflow evaluation. All displayed models use max effort. Refusals count as failures; Astra leads this comparison.
Anthropic’s multimodal evaluation with web tools and code execution. Opus 5.5 uses max effort; other efforts are not specified in this table.
Scientific research in a terminal environment. Astra leads. Reported uncertainty is approximately ±3.5–5 points per model.
Partial-credit computer-use scores at max effort. Opus 5.5 receives credit for completed checkpoints; full-task success is shown separately below.
Chart-reading performance with tools at max effort, from Anthropic’s launch report. Tool-assisted scores should be compared with the same setup.
Large-scale data collection, scored with soft F1. Anthropic uses offline web tools and a 980K task-token budget; compare these rows with one another.
All task checkpoints must pass. These stricter success rates are different from the partial-credit scores above. Source: Opus 5.5 System Card.
Chart recognition without tool assistance. This highlights the model’s direct visual reasoning rather than the higher tool-assisted scores. Source: Anthropic System Card.
These selected settings score 52.8–54.6% on FrontierCode Main. Effort levels differ as shown; costs are Anthropic’s estimates per task.
Anthropic’s standard global API rates. Input tokens cost 20% less than Opus 5. These are token prices, not a SPIRITT subscription price.
Anthropic’s standard global API rates. Output tokens cost 20% less; total workload savings also depend on token usage and effort.
Repeated context costs 60% less to read from cache than with Opus 5. Caching is especially relevant to long coding and agent sessions.
Max effort reaches higher capability but is not the cheapest way to run every task. In AA’s same-index snapshot, Opus 5.5 medium costs $1.34 per task.
Sources: Anthropic’s September 22 release and System Card, Artificial Analysis v4.3.2, and CursorBench. Anthropic’s Opus 5.5 results generally use max adaptive effort, with xhigh on Terminal-Bench 4.0. Safeguard-triggered fallback is included where enabled; Zapier AutomationBench treats refusals as failures. Partial-credit and full-task success rates are labeled separately. API prices are Anthropic’s list prices; tools, special service tiers, and regional options can add charges.
Give the project a workspace, choose the model setup, and work toward a concrete finish line.
Start with your repository, documents, research question, or business goal. Your workspace keeps the browser, files, terminal, integrations, and memory together so your agent can do real work.

Check the Claude models and connection options available in your workspace. Opus 5.5’s medium default is designed for everyday work; raise the effort when a harder problem justifies it.

Ask for the finished outcome, not just an answer: a tested code change, a usable app, or a report with supporting sources. Keep the outputs and decisions together, and review important actions before they go live.

Put your agent in a workspace with the tools, files, and memory to carry a project through.