SPIRITT logoSPIRITTDeepSeekDeepSeek V4 Flash

DeepSeek V4 Flash is cheap agent speed

The July 31, 2026 official 0731 build jumps ~10 AA Intelligence points over the April Flash preview, lands near GPT-5.6 Luna on intelligence at a fraction of the cost, and ships MIT open weights (284B total / 13B active) with a 1M context. On SPIRITT you put that efficiency inside a real agent workspace.

Across X on DeepSeek V4 Flash

Local runs, cache economics, always-use-Flash routing advice, and open-weight hosting vibes.

Official launch card

DeepSeek's official Flash release: same tiny active footprint, much stronger agent post-training, DSpark speculative decoding, Responses API / Codex support.

What DeepSeek built it for

A 284B MoE with only 13B active parameters, 1M context, MIT weights, and first-party API pricing at $0.14 / $0.28 per M tokens with an extreme ~98% cache-hit discount.

0731 is a post-training + agentic coding jump over the April preview: higher Terminal-Bench / DeepSWE / GDPval-AA, lower token use, and explicit coding-agent integrations—not a new pretrain size.

MIT open weights13B active1M context98% cache hitsAgentic codingCodex-ready

AA Intelligence Index

Kimi K3 (max)
57
GPT-5.6 Luna (max)
51
DeepSeek V4 Flash 0731
50
Gemini 3.6 Flash
50
DeepSeek V4 Flash (Apr)
40

AA article: Flash 0731 ~50 (model page max-effort snapshots also show ~52). Near Luna/GLM/Muse 1.1, behind Kimi K3 open-weights lead.

Where V4 Flash is strong vs other models

Same score means nothing alone. These are head-to-head charts so you can see where it leads, where it is close, and what that means for real agent work.

Agentic work
Near frontier

GDPval-AA v2 Elo

Kimi K3 (max)
1687
DeepSeek V4 Flash 0731
1559
GLM-5.2 (max)
1510
V4 Flash (Apr)
1189

Independent AA. 0731 jumps 1189 → 1559; second-highest open-weights once released, behind Kimi K3.

Terminal agents
Near frontier

Terminal-Bench 2.1

Opus-class (ref)
85
DeepSeek V4 Flash 0731 (AA)
79
V4 Flash preview (AA path)
62

AA: 79% (+17 pts vs predecessor). Vendor HF table also lists higher harness figures—prefer AA for cross-lab compare.

Science coding

SciCode (AA)

DeepSeek V4 Flash 0731
50
V4 Flash prior
45

AA: SciCode ~50% (+5 vs predecessor).

Hard exams

Humanity's Last Exam (AA)

DeepSeek V4 Flash 0731
37
V4 Flash prior
32

AA: HLE ~37% (+5 pts).

Speed
Near frontier

Output tokens / sec (AA)

Gemini 3.6 Flash
232
DeepSeek V4 Flash 0731
122
Typical peer median (ref)
70

AA first-party API ~122 tok/s—fast for its intelligence band.

Price
Winner

API input ($ / M tokens)

DeepSeek V4 Flash
$0.14
Gemini 3.6 Flash
$1.50
Kimi K3
$3.00

$0.14 / $0.28 list with ~$0.003 cache hits (~98% discount). Core reason AA cost/task undercuts Luna.

Efficiency
Winner

AA Index output tokens (millions)

DeepSeek V4 Flash 0731
206
V4 Flash prior
234

0731 used ~206M output tokens vs ~234M for prior Flash (−12%) while scoring higher.

Honesty gap

AA-Omniscience hallucination %

Better closed peers (ref)
40%
DeepSeek V4 Flash 0731
84%

Hallucination falls (~96%→84%) but remains high; accuracy flat. Strong agents still need verification loops.

Sources: Artificial Analysis DeepSeek V4 Flash 0731 article and model page; DeepSeek API changelog / HF DeepSeek-V4-Flash-0731 card. AA composite ~50; some max-effort snapshots read ~52. Vendor Terminal-Bench tables can exceed AA harness numbers. On SPIRITT, pair Flash with tools and checks—not blind trust.

How It Works

From zero to DeepSeek V4 Flash running real work in a cloud agent environment

01

Open a workspace

Open a workspace and land in a fully equipped cloud computer: browser, files, terminal, integrations, and memory. No local setup. No thin chat box pretending to be an agent.

Open a SPIRITT workspace
02

Pick DeepSeek V4 Flash

Open the model picker and choose DeepSeek V4 Flash. Same model shipped for agentic work, now inside a workspace that already has tools, browser control, and durable context.

Select DeepSeek V4 Flash in the model picker
03

Build or automate

Tell it what to ship or what to run. V4 Flash can code, call tools, drive the browser, coordinate multi-step work, and keep going while you step away. The point is not another chat window. It is an agentic environment where V4 Flash actually does the job.

Build and automate with DeepSeek V4 Flash

Try DeepSeek V4 Flash in an agentic environment

Cheap, fast, open-weight agent horsepower inside a full cloud workspace.

Questions

Frequently asked questions

01What is DeepSeek V4 Flash?+
DeepSeek's efficient V4 Flash tier (official 0731 build): 284B total / 13B active MoE, 1M context, MIT open weights, and very low API prices with aggressive cache discounts. Built for high-volume agent and coding work.
02How is 0731 different from the April Flash preview?+
Same size and price, much stronger agent post-training: +10 AA Intelligence points, big GDPval-AA and Terminal-Bench jumps, fewer output tokens, Responses API / Codex support, and public MIT weights.
03What is V4 Flash best for?+
Cost-sensitive agent loops, coding assistants, batch tool use, and teams that want open weights or extreme cache economics without paying frontier list prices.
04What are some of the main limitations of this model?+
Text-only I/O on the AA card, still verbose vs medians, hallucination rates remain high even after improvement, and it trails Kimi K3 / closed frontier leaders on raw intelligence.
05What is a good example of using V4 Flash well?+
Run multi-step coding or ops agents in SPIRITT where cache hits dominate cost—repeat large contexts, verify outputs, and keep human checks on factual claims.
06Where is DeepSeek V4 Flash best to test for full agentic capabilities?+
SPIRITT Workspaces. Pick DeepSeek V4 Flash in the model picker and run real work in a fully equipped cloud environment: tools, browser, files, terminal, memory, and durable sessions. That is where the model can act as an agent, not just chat.
Buy from builders who use what they sellBuilt usingSPIRITT