GDPval-AA v2 Elo
Artificial Analysis same-harness result. Qwen is in the leading cluster, but behind the displayed Opus, Grok, and Fable variants.
Qwen built its flagship for sustained coding, research, tool use, and professional work. Artificial Analysis scores it at 58 on the Intelligence Index. SPIRITT serves it through U.S.-operated infrastructure, independently of Qwen's own cloud, with 262K of text context, function calling, and $2/M input plus $6/M output reference pricing.
The official release, blind preference data, hands-on visual work, independent cost evidence, and long-horizon agent reactions.
Meet Qwen3.8-Max, our most capable model to date. A new bar for coding and cowork at 2.4T parameters, including 10+ days of autonomous coding and long-horizon work.
— Qwen (@Alibaba_Qwen) August 3, 2026
Qwen3.8-Max landed at number four on Frontend Code Arena with 1,668 points, one point behind Claude Opus 5 High in the launch snapshot.
— Arena.ai (@arena) August 3, 2026
Qwen3.8-Max beat Fable 5 at building three self-contained 3D physics scenes in this matched test while costing about seven times less.
— atomic.chat (@atomic_chat_hq) August 3, 2026
A visual comparison sent the same sky image to Qwen3.8-Max, Claude Opus 5, Kimi K3, and GPT-5.6 Sol. Qwen was described as honestly on par in that test.
— Ann Nguyen (@ann_nnng) August 3, 2026
During preview, Qwen said the model was improving daily, highlighted a large web-frontend gain, and asked users to keep reporting what breaks.
— Qwen (@Alibaba_Qwen) July 20, 2026
The August 6 snapshot put Qwen3.8-Max at 56 on the Intelligence Index, one point behind Kimi K3 and at a higher cost per task. The live model page has since moved to 58.
— Artificial Analysis (@ArtificialAnlys) August 6, 2026
Artificial Analysis measured more agentic turns and a higher cost per Intelligence task than Qwen3.7-Max, even though per-token prices fell.
— Artificial Analysis (@ArtificialAnlys) August 6, 2026
During preview, Qwen researcher Shuai Bai described the model as a trillion-parameter multimodal system aimed at visual productivity, coding, content creation, and complex task execution.
— Shuai Bai (@shuai_bai_) July 19, 2026
The native model and the exact SPIRITT service are related, but they do not expose the same capability contract.
Qwen3.8-Max is Qwen's 2.4-trillion-parameter flagship, with roughly 95B active parameters and an official focus on coding, research, visual work, and multi-day agents. The model developer's own service exposes the full native 1M multimodal contract.
SPIRITT serves Qwen3.8-Max through U.S.-operated infrastructure, independently of Qwen's own cloud. The current route offers 262K text-only context, function calling, and reference pricing of $2/M input, $0.25/M cached input, and $6/M output. This page keeps that hosted contract explicit instead of borrowing capabilities from a different route.
Artificial Analysis scores Qwen3.8-Max at 58, with about 21 output tokens per second and roughly $0.91 per Intelligence task in the August 29 snapshot. That is a strong challenger profile, not a universal first-place result.
Artificial Analysis model snapshots, August 29, 2026. Qwen3.8-Max scores 58: competitive with the frontier, but behind the current leaders.
Independent rows establish its market position. Qwen's own launch table then shows specific strengths without turning them into an overall winner claim.
Artificial Analysis same-harness result. Qwen is in the leading cluster, but behind the displayed Opus, Grok, and Fable variants.
Artificial Analysis multi-turn banking tool-use evaluation. Winner refers only to the displayed pair.
Qwen launch evaluation. Qwen3.8-Max leads the displayed vendor cohort; this is not an independent rerun.
Qwen launch table. Strong instruction following is one of the clearest first-party wins.
Qwen launch table. Agent setup and environment configuration can materially move OSWorld results.
Artificial Analysis model snapshots, August 29, 2026. Qwen3.8-Max is the slowest displayed stream despite its strong quality score.
Reference list prices for the exact hosted routes discussed on these pages. Lower is better; Max buys more capability at a higher operating price.
Independent sources: Artificial Analysis live Qwen3.8-Max model page and same-harness benchmark views, observed August 29, 2026. Vendor sources: Qwen's August 3 launch report and model catalog. SPIRITT route facts were checked against current hosting documentation and the live model route. Vendor benchmark rows use Qwen's harnesses and settings; run your own repeated evaluation before changing a production route.
From a bounded objective to Qwen3.8-Max working across files, tools, and verification inside a SPIRITT workspace
Start with a repository, research question, or multi-part professional deliverable. State the constraints, source material, and acceptance checks so a long run has a real finish line.

Choose Qwen3.8-Max in the SPIRITT model picker. The SPIRITT route is text-only with 262K context, so bring files and tool results into the workspace rather than assuming the model developer's separate vision contract.

Let the model inspect files, call tools, run commands, and verify intermediate results. Keep human approval on high-impact changes and evidence checks on factual claims.

Run Qwen3.8-Max inside a workspace with files, browser, terminal, tools, memory, and explicit verification.