DeepSeek Guide — whale logoDeepSeek GuideFAN SITE
BENCHMARKS6 MIN READ

DeepSeek V4.1 Flash Benchmarks: Full Official Table

UPDATED: SEP 11, 2026AUTHOR: INDEPENDENT FAN GUIDE
OVERVIEW

Every DeepSeek V4.1 Flash benchmark: Terminal-Bench 2.1 90.6, DeepSWE 74.2, CyberGym 88.1, plus the multi-scaffold table and the effort-cost caveat.

01

How DeepSeek Ran These Benchmarks

All instruct results use maximum reasoning effort (reasoning_effort=100), temperature=1.0, and top_p=0.95. For code-agent benchmarks, the model was evaluated with the Minimal mode of DeepSeek Harness and a 1M-token context window. DeepSWE v1.1 used the mini-SWE harness and SEC-Bench Pro used the Claude Code harness[4].

That methodology matters. The agentic scores are not raw model scores — they are model-plus-harness scores, which is the same caveat that applies to every vendor's agent benchmark. DeepSeek also publishes a separate multi-scaffold table so readers can see how the same model behaves under Claude Code, Codex, OpenCode, and others[4].

Base-model results are evaluated in DeepSeek's internal framework under shared settings, with scores within 0.3 considered equivalent[4].

NOTE

This site has not independently verified any DeepSeek number. Treat all figures as vendor-reported until independent harnesses publish runs[4][5].

02

The Headline Wins: DeepSWE, CyberGym, TB 2.1

On DeepSWE v1.1, V4.1 Flash scores 74.2 — ahead of Claude Opus 5 (74.0) and GPT-5.6 Sol (73.0), and far ahead of its predecessor V4-Pro (62.7) and V4-Flash (54.4). On CyberGym it scores 88.1, the highest in the table. On Terminal-Bench 2.1 it reaches 90.6[4].

The DeepSWE jump is the most striking single number: from 54.4 on V4-Flash to 74.2 on V4.1 Flash in roughly six weeks, and from 62.7 on the V4-Pro flagship. DeepSeek frames this as a consequence of the new architecture plus larger-scale agent-task synthesis in post-training[4].

  • DeepSWE v1.1: 74.2 (Opus 5: 74.0; GPT-5.6 Sol: 73.0)
  • CyberGym: 88.1 (highest listed)
  • Terminal-Bench 2.1: 90.6 (highest listed)
  • AutomationBench: 54.8 (highest listed)
  • Agent's Last Exam: 31.8 (highest listed)
  • Codeforces rating: 3471 (highest listed)[4]
NOTE

These are DeepSeek-run evaluations. TNW and other outlets have not independently verified them[5].

03

Where V4.1 Flash Still Trails

The results are not uniformly dominant. Claude Opus 5 leads V4.1 Flash 43.3 to 30.0 on Terminal-Bench 3.0 and 51.8 to 31.2 on Terminal-Bench 4.0. Opus 5 also leads on HLE (56.3 vs 36.8), ProgramBench (37.0 vs 20.3), and NL2Repo-Bench (75.3 vs 64.0)[4][5].

GPT-5.6 Sol leads GPQA Diamond (94.1 vs 90.9) and SEC-Bench Pro (74.3 vs 62.8). The newer, harder Terminal-Bench versions show the clearest gap: V4.1 Flash wins the older TB 2.1 but loses decisively on TB 3.0 and 4.0, suggesting its strength is concentrated in established agentic tasks rather than frontier ones[4][5].

BenchmarkV4.1 FlashOpus-5.0GPT-5.6 Sol
Terminal-Bench 3.030.043.334.4
Terminal-Bench 4.031.251.839.9
HLE36.856.344.5
ProgramBench20.337.023.0
NL2Repo-Bench64.075.356.8
GPQA Diamond90.993.494.1
SEC-Bench Pro62.874.3
NOTE

DeepSeek's own table includes these losses, which is a point in favor of reading it directly rather than relying on summary graphics[4].

Sponsored
04

The Full Frontier Comparison Table

The official model card compares V4.1 Flash against Claude Opus-5.0, GPT-5.6 Sol, Kimi K3, GLM-5.3, V4-Pro, and V4-Flash across reasoning and agentic categories at max effort[4].

BenchmarkV4.1 FlashOpus-5.0GPT-5.6 SolKimi K3GLM-5.3V4-ProV4-Flash
GPQA Diamond90.993.494.192.988.192.489.9
HLE36.856.344.543.542.042.737.8
Codeforces (Rating)347133483289
MathArena Apex65.665.665.665.358.6
Terminal-Bench 2.190.689.188.888.388.287.982.7
Terminal-Bench 3.030.043.334.417.728.311.87.6
Terminal-Bench 4.031.251.839.912.637.912.47.0
DeepSWE v1.174.274.073.067.566.962.754.4
ProgramBench20.337.023.017.519.015.5
NL2Repo-Bench64.075.356.858.058.061.554.2
CyberGym88.184.580.084.583.376.7
SEC-Bench Pro62.874.356.430.9
ExploitGym15.322.133.715.05.41.8
HLE w/ tools63.963.659.862.560.051.5
AutomationBench54.850.345.846.748.843.237.7
Agents' Last Exam31.828.626.727.628.525.725.2
NOTE

Em dashes mark benchmarks not reported for that model in DeepSeek's official table. HLE figures for V4-Pro, V4-Flash, and GLM are text-only subsets[4].

05

Performance Across Agent Scaffolds

DeepSeek published a scaffold-sensitivity table showing the same model under Claude Code, Codex, OpenCode, Pi, mini-SWE, and three DeepSeek Harness configurations, using N=8 samples on DeepSWE v1.1 and N=3 on Terminal-Bench 2.1[4].

The spread is roughly 9 points on DeepSWE and 6 points on TB 2.1 between the best and worst scaffold. That is a reminder that agentic benchmark headlines are partly harness headlines — a point the Harness-focused community has made repeatedly[4].

ScaffoldDeepSWE v1.1Terminal-Bench 2.1
DeepSeek Harness Minimal74.290.6
mini-SWE agent72.690.3
DeepSeek Harness Standard70.585.8
Claude Code69.888.0
DeepSeek Harness PTC67.685.8
Pi66.286.1
Codex65.684.1
OpenCode65.585.0
NOTE

The official agentic scores use DeepSeek Harness Minimal mode, so they represent the model's ceiling under DeepSeek's own runtime. See the coding agents guide for setup across tools.

06

Base-Model and Multimodal Scores

The base model advances world knowledge and code over V4-Flash-Base but trails V4-Pro-Base on several knowledge and long-context items. Multimodal scores are new to the Flash line[4].

Multimodal base scores: MMMU-Pro 56.5, CVBench 77.9, DocVQA 95.6, and RefCOCO-avg 86.0. V4.1 Flash trained from scratch on a 45T-token multimodal corpus, with sparse attention trained at 64K sequence length and context extended to 1M[4].

BenchmarkV4.1-Flash-BaseV4-Pro-BaseV4-Flash-Base
MMLU-Pro74.173.568.3
AGIEval83.484.483.9
C-Eval92.193.192.1
SimpleQA-Verified42.355.230.1
LongBench-V245.251.544.7
BigCodeBench60.659.256.8
HumanEval79.476.869.5
GSM8K93.092.690.8
MATH61.164.557.4
NOTE

Base-model comparisons use the same internal framework; scores within 0.3 are treated as equivalent by DeepSeek[4].

Sponsored
07

The Reasoning-Effort Cost Caveat

The published table uses maximum effort, which is not the economical setting. DeepSeek's own tests show that increasing effort from 25 to 100 raises DeepSWE v1.1 from 66.0 to 74.2 and Terminal-Bench 2.1 from 82.4 to 90.6 — but consumes roughly 2.5x as many output tokens[5].

DeepSeek says the gains are front-loaded: effort levels between 60 and 80 recover most of the maximum-effort accuracy at less than half the token budget, while the final step to 100 makes agent trajectories 1.6-1.8x longer for comparatively small gains. A team optimizing cost per successful task may find the best operating point away from max[5].

Reasoning effortDeepSWE v1.1Terminal-Bench 2.1Relative output tokens
2566.082.41.0x
60-80most of the gainmost of the gain< 1.25x
100 (leaderboard)74.290.6~2.5x
NOTE

The API exposes presets low/high/max and accepts integers from 1 to 100. See the reasoning effort guide.

08

How to Read Vendor Benchmarks

Every number in this guide comes from DeepSeek's own model card and news post. Independent verification is still pending, and the agentic results depend on the harness, effort level, context limit, and scaffold used[4][5].

The practical takeaway for builders is to test the model on your own workload shape. VentureBeat noted that V4.1 Flash is best suited to input-heavy, repetitive, cacheable workloads that exploit its 8B prefill path and smaller KV cache — and that a model footprint that nearly doubled makes self-hosting harder even as API serving gets cheaper[5].

  • Vendor scores use max effort — your cost/quality point may differ.
  • Agentic scores are harness-dependent; test your own scaffold.
  • Prefer cost per completed task over benchmark position[5]
NOTE

For pricing at each effort level, see the pricing guide.

Sponsored
Sponsored