Terminal-Bench-Science 0.1
Anthropic launch table. Standard error is reported at plus or minus 3.5 to 4.5 points per model.
Anthropic's September release leads the launch-day Artificial Analysis Intelligence Index at 66, reaches 91.4% on Terminal-Bench 2.1, and cuts cache-read pricing by 75%. In SPIRITT, give Fable 5.1 the browser, terminal, files, tools, and persistent context for the hardest multi-step work.
The official launch and independent evidence, followed by six concrete builds: kart racing, a living voxel kingdom, Minecraft, a fantasy Mount Fuji world, Rocket League, and a subway FPS.
We’re introducing Claude Fable 5.1 and Claude Mythos 5.1. They're the world’s most advanced models for coding and knowledge work.
— Claude (@claudeai) September 1, 2026
Fable 5.1 excels at complex, long-running tasks. And its research capabilities offer an early glimpse of how AI models will contribute to scientific progress.
— Claude (@claudeai) September 1, 2026
Claude Fable 5.1 tops the Artificial Analysis Intelligence Index but costs 20% more per task than Fable 5 despite a 75% cache read price cut We supported @AnthropicAI with pre-release evaluation of Claude Fable 5.1. At max effort it scores 66 on the Artificial Analysis Intelligence Index, the highest score we have measured, ahead of Claude Opus 5 (max, 63), Claude Fable 5 (max, 62), GPT-5.6 Sol (max, 61) and Grok 4.6 (high, 61). We evaluated the model with Anthropic's ‘default’ server-side fallback, which routes safety-flagged requests to Claude Opus 4.8 or Claude Opus 5; fallback served ~4% of output tokens across the Intelligence Index. Key takeaways ➤ Frontier Intelligence with improvements across benchmarks: Fable 5.1 gains +4 points on the Intelligence Index over Fable 5. On HLE, Fable 5.1 scores 59.1%, ahead of the previous best of 55.5% from Claude Fable 5. It posts the narrowly highest scores we’ve seen on Terminal-Bench v2.1 (91.4%) and SciCode (62.0%), and on τ³-Banking it gains 9 points over Fable 5 ➤ 75% cache read price cut, but Fable 5.1 still costs more per task: Anthropic has cut the cache read price from $1 to $0.25 per 1M cached input tokens, with standard pricing unchanged at $10/$50 per 1M input/output tokens. Fable 5.1 (max) costs $3.76 per Intelligence Index task, 20% more than Fable 5 (max), because it uses ~1.7x the output tokens. The cache cut saves ~$1.40 per task, concentrated in the agentic evaluations where the majority of input tokens are cache reads. At xhigh effort Fable 5.1 scores 65 at $2.72 per task, $1.04 less than max, but still above Claude Opus 5 (max, 63) at $2.34 ➤ Claude Fable 5.1 holds the upper end of the Intelligence vs Output Tokens per Task Pareto frontier: every model variant scoring higher than GPT-5.6 Sol (medium) on the Intelligence Index is matched or beaten by a Fable 5.1 effort level on both intelligence and token usage ➤ Highest scores on agentic work tasks, but effectively tied with Opus 5: Fable 5.1 sets the highest scores we have measured on GDPval-AA v2 (1,853 Elo, +130 over Fable 5) and AA-Briefcase (1,694 Elo, +122 over Fable 5), our agentic knowledge work evaluations. Against Claude Opus 5 the GDPval-AA v2 lead is within the confidence interval and AA-Briefcase (1,685) is effectively tied, with Fable 5.1 ahead on analytical quality and rubric correctness, but behind on presentation Other model details: ➤ Context window: 1 million tokens, supporting image and text inputs as with Anthropic’s other recent launches ➤ Pricing: Fable 5.1 retains the $10/$50/$12.5 input, output, and cache write prices per million tokens from Fable 5, but cache hits have been reduced to $0.25 per million tokens, a 75% relative reduction from before that will materially reduce agentic workload costs
— Artificial Analysis (@ArtificialAnlys) September 1, 2026
Fable 5.1 is our best model yet for coding, data analysis, computer use, design, presentations, Tag, and the hardest long-running agentic work. This model is a pleasure to work with, and I've been using it for everything.
— Boris Cherny (@bcherny) September 1, 2026
Claude Fable 5.1, the latest in @AnthropicAI's Mythos model class, is now generally available in GitHub Copilot. In our testing it showed strong performance on ➡️ long-running coding tasks ➡️ deep codebase research ➡️ complex agentic workflows Try it out in the GitHub Copilot app, CLI or @code. https://github.blog/changelog/2026-09-01-claude-fable-5-1-generally-available-in-github-copilot/
— GitHub (@github) September 1, 2026
Fable 5.1 from @AnthropicAI on ARC-AGI (Verified): - ARC-AGI-2: 90.0%, $3.12/task - ARC-AGI-1: 97.5%, $1.40/task Across the two benchmarks, Fable 5.1's average cost per task is about 32% lower than Fable 5's, driven by better token efficiency.
— ARC Prize (@arcprize) September 1, 2026
Claude Fable 5.1 ONE SHOTTED this Mario Kart game. This is one of the best results I've had so far and I am SUPER impressed with the game development capabilities. Fable 5.1 is a huge step up from Fable 5.
— BridgeMind (@bridgemindai) September 1, 2026
2.4 million voxels. 1,500 soldiers. 160 working NPCs. one prompt. Fable 5.1 one shotted the largest playable medieval voxel kingdom i’ve ever seen. it includes: • soldiers operating siege equipment • NPCs with jobs who react to the invasion • interiors inside every building • Every building has a function in the city • a fully destructible kingdom • a dragon that breathes fire • the biggest map i’ve produced so far for voxel world generation, Fable 5.1 is the most capable model i’ve tested. not even close. this isn’t a voxel diorama anymore. it’s a living, destructible simulation.
— Knowix (@knowixbuilds) September 1, 2026
Fable 5.1 Minecraft one-shot
— spicylemonade (@spicey_lemonade) September 1, 2026
Need beautiful 3D fantasy or gaming worlds? No problemo. Claude Fable 5.1 one-shot this fantasy ascent to Mt. Fuji entirely in code.
— Techartist (@techartist_) September 1, 2026
Fable 5.1 is NUTS at Game Design. Per usual, RocketLeagueBench is live, and Fable absolutely knocked it out of the park. This is a 2 shot prompt, and I'm certainly impressed. What about you? Super busy today so don't have a ton of time to test right now but here's a few takeaways: - Model is way faster. This took well over an hour for Fable 5. Maybe 20 minutes for fable. - It's MUCH easier to talk to. None of that ridiculous claudese unreadable prose. - It is very good at testing its own work. I checked in on its progress and did not interrupt it. But it was pretty flawed on the first few passes. But it reviewed itself. It basically found many, many bugs on its own, and by the time it finally yielded to me, 98% of the bugs that I saw earlier in the process were gone. This will be the defacto model for the forseeable future for any task that requires 3D design. It will surely excel in CAD. The one thing we don't know yet is how many civilizations it has destroyed :P Anthropic Cooked.
— am.will (@LLMJunky) September 1, 2026
Claude Fable 5.1 Ultracode subway fps game
— BijanBowen (@bijanbowen) September 1, 2026
Fable 5.1 is strongest when the work spans tools, files, decisions, and verification loops. The evidence also shows a real cost and safeguard trade-off.
Claude Fable 5.1 is Anthropic's generally available frontier model for demanding reasoning, long-horizon agentic coding, multistep research, and document, spreadsheet, and presentation work. Official documentation lists a 1M-token context window, 128K maximum output, text and image input, and adaptive thinking that is always on.
Artificial Analysis measured a launch-day Intelligence Index score of 66, the highest in its v4.1.1 snapshot, plus 91.4% on Terminal-Bench 2.1 and 1,853 Elo on GDPval-AA v2. Anthropic's own launch table reports leading results across terminal science, professional work, computer use, automation, and coding.
The $10 input and $50 output price per million tokens did not change from Fable 5. Cache reads fell to $0.25, but Artificial Analysis still measured max effort at roughly 20% more per task because Fable 5.1 used more output tokens. Its tested fallback also served about 4% of output tokens when safeguards intervened.
Claude Fable 5.1 is available through the Fable option in SPIRITT Workspaces. Put it on work where extra judgment, persistence, and verification are worth the premium, then keep consequential approvals and final review with a person.
Artificial Analysis launch snapshot, September 1, 2026. Fable 5.1 max with default fallback led at 66; fallback served about 4% of output tokens across the Index.
The ranking above is independent. The capability cards below use Anthropic's launch table except for the cost card, which uses Artificial Analysis measurements.
Anthropic launch table. Standard error is reported at plus or minus 3.5 to 4.5 points per model.
Anthropic launch table at each model's displayed high-capability setting.
Anthropic launch table. The independent Artificial Analysis launch result also reports 1,853 Elo.
Anthropic used the August 2026 task release. Earlier OSWorld 2.0 numbers are not directly comparable.
Anthropic launch table. The displayed frontier models are separated by less than 1.5 points.
Anthropic launch table. This is a strong relative gain, not broad task reliability.
Anthropic launch table. Cursor separately described Fable 5.1 as its strongest tested model on this benchmark.
Artificial Analysis launch snapshot. Lower is better; the cache cut did not offset higher output-token use at max effort.
Independent source: Artificial Analysis's September 1 Fable 5.1 launch evaluation. Vendor source: Anthropic's Fable 5.1 launch table and system card. Anthropic evaluated Fable 5.1 with production safeguards enabled, and both sources disclose fallback or safeguard effects. Scores can change with model effort, harness, fallback configuration, and benchmark revisions. Test your own highest-value workflows before switching production work.
Choose Claude Fable 5.1 in the picker, then give it a finish line worthy of a frontier model
Open a workspace and land in a fully equipped cloud computer: browser, files, terminal, integrations, and memory. No local setup. No thin chat box pretending to be an agent.

Open the SPIRITT model picker and select Claude Fable 5.1. Use it when long-horizon judgment, difficult debugging, deep research, or polished professional output justifies the premium tier.

Tell it what to ship or what to run. Fable 5.1 can code, call tools, drive the browser, coordinate multi-step work, and keep going while you step away. The point is not another chat window. It is an agentic environment where Fable 5.1 actually does the job.

Open a SPIRITT Workspace, choose Claude Fable 5.1, and give it the evidence, tools, constraints, and acceptance checks for a real end-to-end outcome.