SPIRITT logoSPIRITTZ.aiGLM 5.3 Flash

GLM 5.3 Flash makes agent economics hard to ignore

Previewed as Ox Alpha, Z.ai's 320B MoE activates only 18B parameters per token, adds native vision and a 1M context, and scores 57 on Artificial Analysis at roughly $0.09 per evaluated task. SPIRITT serves this Chinese-developed model through U.S.-operated infrastructure, independently of Z.ai's own cloud, with text, images, tools, and $0.15/M input plus $0.50/M output reference pricing.

GLM 5.3 Flash across X

The official Ox Alpha reveal, independent cost and quality data, coding-agent adoption, architecture notes, and one direct frontier-model caveat.

Flash is the price and compute thesis, not a speed trophy

A newly trained multimodal base cuts active compute while keeping serious coding and agent capability.

What Z.ai changed

GLM 5.3 Flash is a 320B-total, 18B-active MoE with a newly trained base, 45 layers, hybrid sparse and linear attention, Manifold-Constrained Hyper-Connections, and a 30T-token multimodal pretraining corpus. It is released under MIT.

SPIRITT serves this Chinese-developed model through U.S.-operated infrastructure, independently of Z.ai's own cloud, with text and image input, function calling, reasoning, vision, a 1,048,576-token context, and reference pricing of $0.15/M input, $0.03/M cached input, and $0.50/M output.

Artificial Analysis scores it at 57 with roughly 49 output tokens per second and $0.09 per Intelligence task. The value is exceptional; the output stream is not unusually fast for its class.

320B total / 18B activeMIT weights1M contextText + imageHybrid attention$0.15 / $0.50 per M

AA Intelligence Index v4.1.1

Claude Opus 5 max
63
Claude Fable 5 max
62
Grok 4.6 high
61
Qwen3.8-Max
58
GLM 5.3 Flash max
57
Qwen3.8 27B xhigh
52

Artificial Analysis model snapshots, August 29, 2026. GLM 5.3 Flash reaches 57 at a much lower cost than the leading premium models, but it does not lead overall.

The value case, with the frontier gap still visible

Artificial Analysis supplies the independent quality, cost, speed, and knowledge view. Z.ai's table supplies specific coding and tool-use hypotheses.

Independent knowledge work
Near frontier

GDPval-AA v2 Elo

Claude Opus 5 max
1849
Claude Opus 5 xhigh
1817
GLM 5.3 Flash max
1770
Grok 4.6 high
1753
Claude Fable 5 max
1741

Artificial Analysis same-harness result. GLM 5.3 Flash sits in the leading cluster behind the displayed Opus variants.

Vendor terminal agents

Terminal-Bench 2.1

GPT-5.6 Terra
87.4%
Gemini 3.7 Flash
85.8%
Claude Opus 4.8
85%
GLM 5.3 Flash
84.3%
DeepSeek V4 Vision Exp
83.9%

Z.ai launch table. GLM is competitive, but the displayed Terra, Gemini, and Opus runs remain ahead.

Vendor software engineering

DeepSWE v1.1

GPT-5.6 Terra
69.6%
Gemini 3.7 Flash
65.3%
GLM 5.3 Flash
63.4%
DeepSeek V4 Vision Exp
59.3%
Claude Opus 4.8
58%

Z.ai launch table. GLM makes a large within-family jump and beats Opus 4.8 in this setup, while Terra and Gemini lead.

Vendor automation
Near frontier

AutomationBench v1.0.6

Gemini 3.7 Flash
52.3%
GLM 5.3 Flash
48.8%
Claude Opus 4.8
41%
DeepSeek V4 Vision Exp
38.8%
GPT-5.6 Terra
37.2%

Z.ai launch table. GLM ranks second in the displayed cohort and improves 22.6 points over GLM-5.2.

Vendor tool use
Winner

Toolathlon Verified

GLM 5.3 Flash
78.4%
Claude Opus 4.8
76.2%
DeepSeek V4 Vision Exp
75.9%
GPT-5.6 Terra
74.9%

Z.ai obtained results through the official evaluation service and reports pass@1 averaged across three runs.

Independent economics
Winner

Cost per AA Intelligence task

GLM 5.3 Flash max
$0.09
Qwen3.8 27B xhigh
$0.37
Grok 4.6 high
$0.84
Qwen3.8-Max
$0.91

Artificial Analysis, August 29, 2026. Lower is better; GLM 5.3 Flash is the clear value leader in this displayed quality cohort.

Independent throughput

Output tokens / second

Grok 4.6 high
65.5
Qwen3.8 27B xhigh
50.4
GLM 5.3 Flash max
49.4
Qwen3.8-Max
21

Artificial Analysis rolling model snapshots. The word Flash describes efficiency and cost here, not the fastest output stream.

Independent knowledge limit

AA-Omniscience accuracy

GPT-5.6 Terra
47%
GLM 5.3 max
34%
GLM 5.3 Flash
28%

Artificial Analysis. GLM's 28% accuracy is below larger frontier peers even though its hallucination rate is more controlled than some open models.

Reference route price
Winner

Output price ($ / M tokens)

GLM 5.3 Flash
$0.50
Qwen3.8 27B
$3.20
Qwen3.8-Max
$6.00
Kimi K3
$15.00

Reference list prices for the exact hosted routes. Lower is better.

Independent source: Artificial Analysis GLM-5.3-Flash live model page and same-harness evaluation views observed August 29, 2026. Vendor source: Z.ai's August 26 launch report and model-card footnotes. SPIRITT route facts were checked against current hosting documentation and the live model route. Flash offers exceptional price-performance, but its speed, knowledge accuracy, and large self-hosting footprint remain real tradeoffs.

How It Works

From a multimodal brief to GLM 5.3 Flash seeing, coding, using tools, and refining the result inside one workspace

01

Open the full working context

Bring the repository, screenshots, documents, spreadsheets, or research packet into a SPIRITT workspace. The SPIRITT route can use text and images across a 1M context.

Open a SPIRITT workspace
02

Pick GLM 5.3 Flash

Choose GLM 5.3 Flash in the model picker. Use its low-cost reasoning and function calling for iterative work, while keeping the expected output and evidence criteria explicit.

Select GLM 5.3 Flash in the model picker
03

Let vision close the coding loop

Have it build, render, inspect, and refine. Keep tests and screenshots as the oracle, and route unusually difficult or high-risk decisions to stronger review rather than relying on price alone.

Build and automate with GLM 5.3 Flash

Spend frontier-model money only where the task earns it

GLM 5.3 Flash brings a 1M multimodal route, tools, files, terminal, browser, and memory into one low-cost agent workspace.

Questions

GLM 5.3 Flash FAQ

01What is GLM 5.3 Flash?+
GLM 5.3 Flash is Z.ai's August 2026 open-weight multimodal model with 320B total parameters, 18B active parameters, hybrid sparse and linear attention, a 1M context, and MIT licensing.
02How is GLM 5.3 Flash different from GLM 5.3 and GLM 5.2?+
Flash starts from a new 320B/18B multimodal base built for efficient serving. Full GLM 5.3 is a larger text-focused flagship. Z.ai reports large Flash gains over GLM 5.2 on DeepSWE and AutomationBench at roughly one-tenth the list price.
03Where does SPIRITT host GLM 5.3 Flash?+
SPIRITT serves this Chinese-developed model through U.S.-operated infrastructure, independently of Z.ai's own cloud. This describes the hosting operator and does not promise U.S.-only data residency.
04What is GLM 5.3 Flash best for?+
Cost-sensitive coding agents, multimodal frontend and document work, tool-heavy automation, and long-context workflows where price per accepted task matters more than maximum token streaming speed.
05What are the main limitations of GLM 5.3 Flash?+
Artificial Analysis measures about 49 output tokens per second and 28% knowledge accuracy on AA-Omniscience. The model is somewhat verbose, and the open checkpoint is still hundreds of gigabytes before runtime and context memory.
06What are the reference token economics for GLM 5.3 Flash?+
The hosted route's reference pricing is $0.15 per million input tokens, $0.03 per million cached input tokens, and $0.50 per million output tokens. SPIRITT plan pricing and usage accounting are separate product terms.
07Can I connect a Z.ai subscription from SPIRITT Subscriptions?+
Not currently. Choose GLM 5.3 Flash through SPIRITT Models. The Subscriptions tab currently connects ChatGPT, Claude, and Grok accounts, not Z.ai accounts.
08What is a good example of using GLM 5.3 Flash well?+
Ask it to turn a product brief and reference screenshots into a working responsive page: infer the design system, implement the code, launch it, compare desktop and mobile screenshots, fix visual differences, and return the final build evidence.
09Where is GLM 5.3 Flash best to test for full agentic capabilities?+
SPIRITT Workspaces. Pick GLM 5.3 Flash in the model picker and run real work in a fully equipped cloud environment: tools, browser, files, terminal, memory, and durable sessions. That is where the model can act as an agent, not just chat.
Buy from builders who use what they sellBuilt usingSPIRITT