DeepSeek V4 Pro + DeepSeek Harness: The Official Agent Stack
Pair DeepSeek V4 Pro 0813 with the official Harness (dsh) agent runtime: npx @deepseek-ai/dsh web, default model routing, official benchmark stack.
Why the Official Combo
DeepSeek's own agent benchmarks are measured on this exact pair: V4 Pro 0813 running inside DeepSeek Harness minimal mode at max reasoning effort, temperature 1.0, top_p 0.95[1][5]. Terminal-Bench 2.1 hits 87.9, Cybergym 83.3 (rank #1), AutomationBench 31.8 (rank #1)[1].
The Harness (dsh, v0.1 developer preview, MIT) is the agent runtime where every capability — models, tools, skills, sessions, sandboxes, storage, loops — is a plugin[6]. V4 Pro is the flagship model that plugs in by default.
The result is the most 'official' agent stack in the DeepSeek ecosystem: the same code path DeepSeek uses to publish its own agent scores[1][5].
Install & First Run
One command launches the local web UI; the CLI exposes headless and profile modes for scripting[6][7].
- Default web UI listens on 127.0.0.1:3080[7].
- First run walks through onboarding: workspace, model, preset, session controls[7].
- CLI profiles (headless, acp, etc.) for scripting and agent-client integration[6].
# quick start (web UI on http://127.0.0.1:3080)
npx @deepseek-ai/dsh web
# or build from source
git clone https://github.com/deepseek-ai/deepseek-harness.git
cd deepseek-harness
pnpm install && pnpm run build
pnpm dsh web
# headless one-shot task
dsh --profile headless "Summarize this repo"
# Node requirement: ^22.19.0 or >= 24.0.0Default Model & API Key
dsh routes model calls through the DeepSeek API by default — you provide a DeepSeek API key during onboarding, and the default model selection is V4 Pro (with Flash as the lighter option)[7]. Because models are plugins, you can swap providers in configuration without touching source[6].
Model calls are billed at DeepSeek API rates ($0.435/$0.87 per 1M for Pro before the 8/16 change; $0.66/$1.98 off-peak after)[4]. KV-cache behavior flows through: stable harness prompts reuse prefix caches, cutting input cost substantially[4].
Because the model is a plugin, you can also point dsh at OpenRouter or a self-hosted vLLM endpoint without touching the harness source — the configuration screen accepts any OpenAI-compatible provider[6][7]. That makes the combo portable: the same session files and plugin set run against V4 Pro today and against a different provider tomorrow.
# During onboarding you supply:
# DEEPSEEK_API_KEY=sk-... (platform.deepseek.com)
# Default model: deepseek-v4-pro (0813 GA)
# Lighter default: deepseek-v4-flash (0731)
# Swap model in config (provider/plugin level):
# models → deepseek provider → model: deepseek-v4-flashRuntime Modes for Pro Workloads
dsh ships four presets. For V4 Pro's strong reasoning, Standard and Code modes are the daily drivers; Minimal mode is exactly what DeepSeek uses for benchmarking[6].
Every run is traceable: append-only session logs capture prompts, reasoning, tool calls, and subagent scheduling, viewable in the Trajectory view — useful for debugging long Pro agent runs and for reproducing cost spikes[6].
| Mode | What it gives you | Best for V4 Pro |
|---|---|---|
| Standard | Full toolset: shell, file edits, search, skills, subagents | Daily agent work[6] |
| Code | Standard + TypeScript program orchestration | Complex multi-step workflows[6] |
| Minimal | bash + str_replace_editor only | Benchmarking / fair model comparison[6] |
| Creator | Inspect runtime, test Cordis plugins live | Building custom harnesses[6] |
Mode descriptions from DeepSeek's official harness page[6].
Benchmarks, Caveats & Alternatives
Official scores use this stack at max effort; independent harnesses measure lower — Vals' reference-harness Terminal-Bench run put Pro at 54.68% versus the official 87.9[3]. Treat official agent numbers as the ceiling, not the floor.
If you want the official stack today, accept the preview-stage churn. If you want stability, OpenCode + V4 Pro gets you most of the way with none of the new-runtime risk — and you can revisit dsh when the plugin ecosystem matures[2][6].
- Official (Harness minimal + max): TB2.1 87.9, Cybergym 83.3 #1[1].
- Independent (reference harness): TB2.1 ~54.7 — the harness matters as much as the model[3].
- Alternative: OpenCode + V4 Pro — lighter, config-only, 96.40% SWE-bench Verified (neutral harness)[2].
- dsh is developer preview: breaking changes expected across 0.1.x releases[6].