GDPval-AA v2 Elo
Independent AA. 0731 jumps 1189 → 1559; second-highest open-weights once released, behind Kimi K3.
DeepSeek V4 FlashThe July 31, 2026 official 0731 build jumps ~10 AA Intelligence points over the April Flash preview, lands near GPT-5.6 Luna on intelligence at a fraction of the cost, and ships MIT open weights (284B total / 13B active) with a 1M context. On SPIRITT you put that efficiency inside a real agent workspace.
Local runs, cache economics, always-use-Flash routing advice, and open-weight hosting vibes.
Been using Deepseek Flash V4 0731 for a week now locally on 4x RTX Pro 6000 and got used to the speed so much that GPT-5.6-Sol feels unbearablly slow. 0731 is my favorite model now and been able to run it locally is the gift from the gods.
— ༽Europurr༼ (@vrloom) August 9, 2026
I've had alot of success using deepseek V4 Flash with Luna subagents with a set goal. This combo for some reason is working surprisingly well. Could probably switch them around too but haven't tested that yet.
— m͓̽o͓̽u͓̽s͓̽e (@dunder_bird) August 9, 2026
Yo currently debugging deepseek v4 0731 flash locally
— Hemant (@heman10x) August 9, 2026
Ran an A/B on DeepSeek-V4-Flash-0731 across 2× DGX Spark (GB10), TP=2. Same model weights. Same nodes. Same prompt set. Only the serve config changed. Baseline (what I had) vs @Tech2Wild / tonyd2wild DSpark+NVFP4 recipe. Spoiler: concurrency is where it stops being subtle.
— Steve Goldsby (@GoldsbySteve) August 9, 2026
Just finished a busy weekend running deepseek v4 flash 0731 on both single and dual dgx nodes - thanks for your recipes! They are great!
— Jacopo Nardiello (@jnardiello) August 9, 2026
Meta drops Muse Code(coding agent) GitHub launches stacked PRs DeepSeek makes v4-flash-0731 default for all users. Qwen launches qwen 3.8 OpenAI slashes prices What else am I missing?
— Noah Sheldon (@noah__sheldon) August 9, 2026
以防你不知道opencode go有多离谱,5刀(首月半价)可以换出来103.59刀的deepseek,按照97.56%的缓存命中率,一个月可以用出来131.8亿deepseek v4 flash的token,换言之是deepseek官方api的0.5折不到🤣请问这和白送有啥区别?
— Zesen Huang (@zesenhuang) August 9, 2026
DeepSeek V4 Flash 0731, SemiAnalysis’e göre Nemotron 3 Ultra’yı agent görevlerinde geçerken 4,2 kat daha az aktif parametre kullanıyor. Resmi Terminal Bench 2.1 skoru 82,7. AI yarışında verimlilik de önemli.
— Salih Hayli (@salihhayli) August 9, 2026
DeepSeek V4 Flash 0731 is now open weights under MIT, among top open-weights models on the Intelligence Index.
— Artificial Analysis (@ArtificialAnlys) July 31, 2026
DeepSeek's official Flash release: same tiny active footprint, much stronger agent post-training, DSpark speculative decoding, Responses API / Codex support.
A 284B MoE with only 13B active parameters, 1M context, MIT weights, and first-party API pricing at $0.14 / $0.28 per M tokens with an extreme ~98% cache-hit discount.
0731 is a post-training + agentic coding jump over the April preview: higher Terminal-Bench / DeepSWE / GDPval-AA, lower token use, and explicit coding-agent integrations—not a new pretrain size.
AA article: Flash 0731 ~50 (model page max-effort snapshots also show ~52). Near Luna/GLM/Muse 1.1, behind Kimi K3 open-weights lead.
Same score means nothing alone. These are head-to-head charts so you can see where it leads, where it is close, and what that means for real agent work.
Independent AA. 0731 jumps 1189 → 1559; second-highest open-weights once released, behind Kimi K3.
AA: 79% (+17 pts vs predecessor). Vendor HF table also lists higher harness figures—prefer AA for cross-lab compare.
AA: SciCode ~50% (+5 vs predecessor).
AA: HLE ~37% (+5 pts).
AA first-party API ~122 tok/s—fast for its intelligence band.
$0.14 / $0.28 list with ~$0.003 cache hits (~98% discount). Core reason AA cost/task undercuts Luna.
0731 used ~206M output tokens vs ~234M for prior Flash (−12%) while scoring higher.
Hallucination falls (~96%→84%) but remains high; accuracy flat. Strong agents still need verification loops.
Sources: Artificial Analysis DeepSeek V4 Flash 0731 article and model page; DeepSeek API changelog / HF DeepSeek-V4-Flash-0731 card. AA composite ~50; some max-effort snapshots read ~52. Vendor Terminal-Bench tables can exceed AA harness numbers. On SPIRITT, pair Flash with tools and checks—not blind trust.
From zero to DeepSeek V4 Flash running real work in a cloud agent environment
Open a workspace and land in a fully equipped cloud computer: browser, files, terminal, integrations, and memory. No local setup. No thin chat box pretending to be an agent.

Open the model picker and choose DeepSeek V4 Flash. Same model shipped for agentic work, now inside a workspace that already has tools, browser control, and durable context.

Tell it what to ship or what to run. V4 Flash can code, call tools, drive the browser, coordinate multi-step work, and keep going while you step away. The point is not another chat window. It is an agentic environment where V4 Flash actually does the job.

Cheap, fast, open-weight agent horsepower inside a full cloud workspace.