DeepSeek Guide — whale logoDeepSeek GuideFAN SITE
MODEL GUIDE8 MIN READ

What Is DeepSeek V4 Flash? Full Guide to the 0731 Release

UPDATED: AUG 1, 2026AUTHOR: INDEPENDENT FAN GUIDE
OVERVIEW

DeepSeek V4 Flash is DeepSeek's 284B-parameter open-weight model with 1M context. Here's the full guide to the July 31, 2026 official release.

01

What Is DeepSeek V4 Flash?

DeepSeek V4 Flash is DeepSeek's fast, efficient, and economical tier of the V4 family. It is a mixture-of-experts (MoE) model with 284B total parameters and 13B active parameters per token, a 1M-token context window, and a maximum output of 384K tokens[3].

Flash debuted on April 24, 2026[4] alongside the flagship V4 Pro. It is text-only, ships with MIT-licensed open weights, and is designed for high-volume, low-latency work. DeepSeek says Flash reasoning closely approaches V4 Pro, and Flash performs on par with V4 Pro on simple agent tasks.

The model card is still the best single source for the architecture: hybrid attention built from Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA), with FP4 + FP8 mixed precision. The 1M context window makes it a practical choice for million-token document and code tasks.

  • 284B total / 13B active MoE parameters
  • 1M context window (1,048,576 tokens); max output 384K (393,216 tokens)
  • Text-only input with hybrid CSA + HCA attention
  • FP4 + FP8 mixed precision
  • MIT-licensed open weights on Hugging Face
SpecDeepSeek V4 Flash
Total parameters284B (MoE)
Active parameters13B
Context window1M (1,048,576 tokens)
Max output384K (393,216 tokens)
Input modalityText only
PrecisionFP4 + FP8 mixed
AttentionCompressed Sparse Attention (CSA) + Heavily Compressed Attention (HCA)
LicenseMIT
02

What the 0731 Official Release Actually Is

On July 31, 2026, DeepSeek moved the Flash API to official release (public beta) under build number DeepSeek-V4-Flash-0731[1]. The key fact: this is not a new model. The official changelog states that 0731 keeps the same model architecture and size as DeepSeek-V4-Flash-Preview, and was only re-post-trained.

Because the architecture and parameter count are unchanged, context length, pricing, and the model slug all stay the same. Setting model=deepseek-v4-flash automatically serves the latest build, so most existing API code needs no changes at all.

The upgrade is limited to the Flash API. The V4-Pro API and the APP/WEB models are unchanged, and DeepSeek says the official release of V4-Pro will follow soon.

  • Re-post-training focused on agentic and coding capabilities
  • Native support for the Responses API format
  • Dedicated Codex integration support[5]
  • Release channel moved from preview to public beta
NOTE

If you were already calling model=deepseek-v4-flash, the July 31 upgrade took effect automatically. No code change is required to get the 0731 build.

03

Flash Preview vs 0731: Exact Differences

The differences between the preview and the 0731 build are easy to overstate. The model itself is unchanged: same 284B/13B architecture, same 1M context, same price. What changed is training and tooling, not the model.

Two legacy API aliases, deepseek-chat and deepseek-reasoner, retired on July 24, 2026 at 15:59 UTC. Migration is a one-line change: set model to deepseek-v4-flash (or deepseek-v4-pro); the base_url and API key stay the same. Note that deepseek-reasoner mapped to Flash thinking mode, not Pro.

AspectFlash Preview (Apr 2026)Flash 0731 (Jul 31, 2026)
Architecture284B / 13B MoESame architecture, same size
Post-trainingOriginalRe-post-trained (agentic focus)
Context / output1M in / 384K outUnchanged
Price (input / output)$0.14 / $0.28Unchanged
Responses APINot supportedNative support
Codex integrationNoneDedicated one-click setup script
Release channelPreviewPublic beta (official)
04

Agent Benchmarks: Flash 0731 vs V4-Pro-Preview

DeepSeek reports that Flash 0731 wins all nine published agent benchmarks[8] against its own V4-Pro-Preview. The gains are most dramatic in software engineering, with DeepSWE jumping from 7.3 on the preview to 54.4 — roughly a 645% increase.

DeepSeek says Flash 0731 is broadly competitive with the strongest proprietary models available. In one external comparison, Terminal-Bench 2.1 at 82.7 also beats Kimi K3's 76.1 from about two weeks earlier.

All nine figures are DeepSeek's self-reported vendor numbers. Independent third parties cannot re-run the harness yet, so treat them as directional until the evaluation harness is open-sourced.

BenchmarkFlash 0731V4-Pro-PreviewFlash Preview
Terminal-Bench 2.182.772.161.8
DeepSWE54.47.3
Cybergym76.7
NL2Repo54.2
Toolathlon (verified)70.3
Agent Last Exam25.2
Automation Bench (Public)25.1
DSBench-FullStack68.7
DSBench-Hard59.6
NOTE

Figures marked with an em dash were not reported in the official 0731 model card for that model. All scores above are vendor-reported by DeepSeek.

05

Pricing: $0.14 Input / $0.28 Output

The 0731 release keeps the preview price. Input costs $0.14 per 1M tokens on a cache miss and output costs $0.28 per 1M tokens. Cached input is dramatically cheaper at $0.0028 per 1M tokens — roughly a 98% discount[2].

Flash output is about a third of V4 Pro's output price, which is why heavy or high-volume agent workloads default to Flash. New accounts also start with roughly 5 million free tokens, valid for about 30 days.

DeepSeek plans peak/off-peak pricing: peak hours (Beijing time UTC+8, daily 9:00-12:00 and 14:00-18:00) will cost 2x the base price. The effective date has not been announced.

  • Cache hits are ~98% cheaper — reuse prompts with stable prefixes to maximize your cache hit rate.
  • Flash supports 2,500 concurrent requests, five times V4 Pro's 500.
  • For high-volume, low-latency work, Flash is the default; reserve Pro for heavy reasoning.
Per 1M tokensV4 FlashV4 Pro (reference)
Input (cache miss)$0.14$0.435
Input (cache hit)$0.0028$0.003625
Output$0.28$0.87
Concurrency limit2,500500
06

DSpark: The Speed Module in 0731 Weights

The 0731 weights ship with DSpark, DeepSeek's speculative decoding framework[7]. Short for Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation, DSpark was open-sourced under MIT on June 27, 2026 and is co-authored by DeepSeek founder Liang Wenfeng and a Peking University team.

In production, DeepSeek reports per-user generation speeds 60-85% faster for V4 Flash than the older MTP-1 baseline, with no weight changes and no extra hardware.

DSpark is also why the Hugging Face repo displays 304B parameters: the extra count is the draft module, not a larger base model. The base model remains 284B.

NOTE

DSpark accelerates per-user decoding without changing weights or adding hardware. DeepSeek also measured 51% higher total throughput on Flash in its own tests.

example_code.py
vllm serve deepseek-ai/DeepSeek-V4-Flash-0731 --trust-remote-code \
  --kv-cache-dtype fp8 --block-size 256 --data-parallel-size 4 \
  --enable-expert-parallel --moe-backend deep_gemm_mega_moe \
  --attention-config '{"use_fp4_indexer_cache": true}' \
  --speculative-config '{"method":"dspark","num_speculative_tokens":7,"draft_sample_method":"greedy"}'
07

How to Access DeepSeek V4 Flash

There are two official ways to use V4 Flash: the API and the chat.deepseek.com web app. On the web, Instant Mode runs the fast Flash tier for quick, low-cost responses; Expert Mode targets heavier reasoning.

For the API, keep the base_url you already use: https://api.deepseek.com for OpenAI-compatible calls, or https://api.deepseek.com/anthropic for the Anthropic format. Set model to deepseek-v4-flash, and the latest official build is served automatically.

Thinking is on by default[6]. In multi-turn agent loops with tool calls, you must pass the previous turn's reasoning_content back with the message, or the reasoning chain breaks. reasoning_effort supports three levels: low, high, and max. DeepSeek recommends max output of at least 384K tokens for high and max effort, with temperature 1.0 and top_p 1.0 (0.95 in agentic settings).

The 0731 build adds two developer-facing capabilities: native Responses API format support (currently Flash only, with Pro expected in early August 2026) and a dedicated Codex integration.

  • OpenAI ChatCompletions and Anthropic-compatible endpoints
  • JSON Output and Tool Calls
  • Responses API (Flash only today)
  • Chat Prefix Completion (Beta) and FIM Completion (non-thinking mode only)
  • Chain-of-thought returned via reasoning_content
NOTE

Codex users: run the official one-click setup — bash <(curl -fsSL https://cdn.deepseek.com/api-docs/codex-deepseek-setup-en.sh) on macOS/Linux, or irm https://cdn.deepseek.com/api-docs/codex-deepseek-setup-en.ps1 | iex on Windows.

example_code.py
curl https://api.deepseek.com/chat/completions \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deepseek-v4-flash",
    "messages": [{"role": "user", "content": "Explain mixture of experts in one paragraph"}],
    "thinking": {"type": "enabled"},
    "reasoning_effort": "high"
  }'
08

V4 Flash vs V4 Pro: Where Each Fits

DeepSeek V4 Flash is the economical tier of the V4 family, and V4 Pro is the flagship. Flash handles fast, high-volume, agent-heavy work; Pro brings the biggest model and the strongest reasoning at a higher price.

DeepSeek's official positioning is that Flash reasoning closely approaches V4 Pro, and Flash is on par with V4 Pro on simple agent tasks. After 0731, Flash now beats V4-Pro-Preview on all nine published agent benchmarks — while costing roughly a third as much for output.

V4 Pro's official release will follow soon, according to DeepSeek. Until then, the V4-Pro API continues to serve the preview build, unchanged by the Flash update.

V4 FlashV4 Pro
Parameters284B total / 13B active1.6T total / 49B active
Context window1M1M
Input (cache miss)$0.14 / 1M$0.435 / 1M
Output$0.28 / 1M$0.87 / 1M
Concurrency2,500500
Best forFast, cheap agent tasksHeavy reasoning, flagship work