GDPval-AA v2 Elo
Artificial Analysis same-harness result. GLM 5.3 Flash sits in the leading cluster behind the displayed Opus variants.
Previewed as Ox Alpha, Z.ai's 320B MoE activates only 18B parameters per token, adds native vision and a 1M context, and scores 57 on Artificial Analysis at roughly $0.09 per evaluated task. SPIRITT serves this Chinese-developed model through U.S.-operated infrastructure, independently of Z.ai's own cloud, with text, images, tools, and $0.15/M input plus $0.50/M output reference pricing.
The official Ox Alpha reveal, independent cost and quality data, coding-agent adoption, architecture notes, and one direct frontier-model caveat.
Introducing GLM-5.3-Flash: native multimodality, a 1M-token context, 320B total and 18B active parameters, MIT weights, and the reveal that it was previewed as Ox Alpha.
— Z.ai (@Zai_org) August 26, 2026
GLM-5.3 Flash, formerly Ox Alpha, became available on OpenCode Go with a temporary double-usage launch offer.
— OpenCode (@opencode) August 26, 2026
GLM-5.3-Flash scored 57 on the Intelligence Index at roughly $0.09 per task, landing on the Intelligence versus cost-per-task Pareto frontier in that snapshot.
— Artificial Analysis (@ArtificialAnlys) August 26, 2026
Cline said GLM-5.3 Flash was its fastest-growing model, reaching more than 11% of traffic in less than a week during the Ox Alpha launch period.
— Cline (@cline) August 26, 2026
The release was highlighted as a newly trained efficient base, not a distillation of GLM-5.3, with 320B total and 18B active parameters.
— David Hendrickson (@TeksEdge) August 26, 2026
The hybrid attention design, Chinese-chip inference capacity, and visual computer-use loop stood out in a launch-day architecture reaction.
— Chinmay (@ChinmayKak) August 26, 2026
GLM 5.3 Flash is good, but the issues become visible quickly if you are used to frontier models.
— Bridgebench (@bridgebench) August 28, 2026
The launch was summarized as the first natively multimodal GLM-5 model, close to Opus 4.8 on coding and agentic tasks at roughly one-tenth the provider price.
— alex getman (@alexgetmancom) August 26, 2026
A newly trained multimodal base cuts active compute while keeping serious coding and agent capability.
GLM 5.3 Flash is a 320B-total, 18B-active MoE with a newly trained base, 45 layers, hybrid sparse and linear attention, Manifold-Constrained Hyper-Connections, and a 30T-token multimodal pretraining corpus. It is released under MIT.
SPIRITT serves this Chinese-developed model through U.S.-operated infrastructure, independently of Z.ai's own cloud, with text and image input, function calling, reasoning, vision, a 1,048,576-token context, and reference pricing of $0.15/M input, $0.03/M cached input, and $0.50/M output.
Artificial Analysis scores it at 57 with roughly 49 output tokens per second and $0.09 per Intelligence task. The value is exceptional; the output stream is not unusually fast for its class.
Artificial Analysis model snapshots, August 29, 2026. GLM 5.3 Flash reaches 57 at a much lower cost than the leading premium models, but it does not lead overall.
Artificial Analysis supplies the independent quality, cost, speed, and knowledge view. Z.ai's table supplies specific coding and tool-use hypotheses.
Artificial Analysis same-harness result. GLM 5.3 Flash sits in the leading cluster behind the displayed Opus variants.
Z.ai launch table. GLM is competitive, but the displayed Terra, Gemini, and Opus runs remain ahead.
Z.ai launch table. GLM makes a large within-family jump and beats Opus 4.8 in this setup, while Terra and Gemini lead.
Z.ai launch table. GLM ranks second in the displayed cohort and improves 22.6 points over GLM-5.2.
Z.ai obtained results through the official evaluation service and reports pass@1 averaged across three runs.
Artificial Analysis, August 29, 2026. Lower is better; GLM 5.3 Flash is the clear value leader in this displayed quality cohort.
Artificial Analysis rolling model snapshots. The word Flash describes efficiency and cost here, not the fastest output stream.
Artificial Analysis. GLM's 28% accuracy is below larger frontier peers even though its hallucination rate is more controlled than some open models.
Reference list prices for the exact hosted routes. Lower is better.
Independent source: Artificial Analysis GLM-5.3-Flash live model page and same-harness evaluation views observed August 29, 2026. Vendor source: Z.ai's August 26 launch report and model-card footnotes. SPIRITT route facts were checked against current hosting documentation and the live model route. Flash offers exceptional price-performance, but its speed, knowledge accuracy, and large self-hosting footprint remain real tradeoffs.
From a multimodal brief to GLM 5.3 Flash seeing, coding, using tools, and refining the result inside one workspace
Bring the repository, screenshots, documents, spreadsheets, or research packet into a SPIRITT workspace. The SPIRITT route can use text and images across a 1M context.

Choose GLM 5.3 Flash in the model picker. Use its low-cost reasoning and function calling for iterative work, while keeping the expected output and evidence criteria explicit.

Have it build, render, inspect, and refine. Keep tests and screenshots as the oracle, and route unusually difficult or high-risk decisions to stronger review rather than relying on price alone.

GLM 5.3 Flash brings a 1M multimodal route, tools, files, terminal, browser, and memory into one low-cost agent workspace.