DeepSeek V4 Flash API Setup: Base URL, Models & Your First Call
DeepSeek V4 Flash API setup in minutes: create your API key, set the base URL, and make your first call with curl, Python, or Node.js.
Create Your API Key
Everything starts with a DeepSeek API key. Open platform.deepseek.com/api_keys, sign in, and create a new secret key. Copy it immediately — the platform shows it only once.[8]
The official DeepSeek API docs open their quickstart with the same instruction: create an API key first.[1] It is the single prerequisite for every request in this guide, and it is the only credential you pass in the Authorization header.
- Go to platform.deepseek.com/api_keys, the official key creation page.
- Create a new secret key and copy it right away. It is displayed only once.
- Store it in an environment variable named DEEPSEEK_API_KEY.
- Never commit the key to git. Keep it out of your repository entirely.
Keys follow the sk- prefix. If you lose one, you cannot recover it — create a new key and rotate the old one.
Set the Correct Base URL
The DeepSeek V4 Flash API is 100% OpenAI-compatible, and it also speaks the Anthropic format. The OpenAI-compatible base URL is https://api.deepseek.com. The Anthropic-compatible endpoint is https://api.deepseek.com/anthropic.
The official changelog is explicit: keep the base URL unchanged and only update the model.[2] DeepSeek does not version its endpoint per release, so nothing changes here when the model gets upgraded — including the July 31, 2026 public-beta release.
Clients that already target OpenAI or Anthropic keep their existing SDKs. You change only the endpoint and the model name.
| API Format | Base URL |
|---|---|
| OpenAI-compatible | https://api.deepseek.com |
| Anthropic-compatible | https://api.deepseek.com/anthropic |
Both endpoints serve the same model. Pick the format that matches the SDK or agent you already use.
// OpenAI-compatible client (Python)
client = OpenAI(
api_key=os.environ.get("DEEPSEEK_API_KEY"),
base_url="https://api.deepseek.com")
// Anthropic-compatible client (env config)
export ANTHROPIC_BASE_URL=https://api.deepseek.com/anthropicPick the Model: deepseek-v4-flash
Set the model parameter to deepseek-v4-flash. This slug now points to DeepSeek-V4-Flash-0731, the official public-beta release from July 31, 2026.
The legacy names deepseek-chat and deepseek-reasoner were retired on July 24, 2026. Any code still sending them now gets an error. New integrations should use deepseek-v4-flash, or deepseek-v4-pro for the flagship model.
deepseek-v4-flash is a 284B-parameter Mixture-of-Experts model with 13B active parameters, a 1M-token context window, and a 384K-token max output.[3]
- JSON Output: yes
- Tool Calls: yes
- Responses API: yes (Flash only — Pro expected early August 2026)
- Anthropic API: yes
- FIM completion: yes (non-thinking mode only)
| Spec | Value |
|---|---|
| Model name | deepseek-v4-flash |
| Current version | DeepSeek-V4-Flash-0731 |
| Parameters | 284B total / 13B active (MoE) |
| Context window | 1M tokens (1,048,576) |
| Max output | 384K tokens |
| Input price (cache miss) | $0.14 / 1M tokens |
| Input price (cache hit) | $0.0028 / 1M tokens |
| Output price | $0.28 / 1M tokens |
| Concurrency limit | 2,500 concurrent requests (account-level) |
If you are migrating, the change is one line: replace deepseek-chat or deepseek-reasoner with deepseek-v4-flash. The base URL and API key stay the same.
Make Your First Call with cURL
A single curl request is the fastest way to prove your key works. This call hits the Chat Completions endpoint with the thinking parameters from DeepSeek's official quickstart.
A successful response returns the assistant text under choices[0].message.content. The thinking and reasoning_effort fields are optional — thinking is on by default — but sending them explicitly makes your intent clear.
If the call fails, read the HTTP status code and match it against the error table in Step 8.[4]
Export DEEPSEEK_API_KEY in your shell first, and let it expand in the header. Do not paste the raw key into the command.
curl https://api.deepseek.com/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer ${DEEPSEEK_API_KEY}" \
-d '{
"model": "deepseek-v4-flash",
"messages": [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Hello!"}
],
"thinking": {"type": "enabled"},
"reasoning_effort": "high",
"stream": false
}'First Call in Python & Node.js
Because the DeepSeek V4 Flash API is OpenAI-compatible, the standard openai SDK works with a one-line change: set base_url to https://api.deepseek.com and send model deepseek-v4-flash.
In the Python SDK, the thinking flag goes inside extra_body because the official SDK does not type it as a first-class parameter. The Node.js SDK accepts thinking directly.
# pip3 install openai
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ.get('DEEPSEEK_API_KEY'),
base_url="https://api.deepseek.com")
response = client.chat.completions.create(
model="deepseek-v4-flash",
messages=[
{"role": "system", "content": "You are a helpful assistant"},
{"role": "user", "content": "Hello"},
],
stream=False,
reasoning_effort="high",
extra_body={"thinking": {"type": "enabled"}})
print(response.choices[0].message.content)
// --- Node.js (OpenAI SDK) ---
import OpenAI from "openai";
const openai = new OpenAI({
baseURL: 'https://api.deepseek.com',
apiKey: process.env.DEEPSEEK_API_KEY,
});
async function main() {
const completion = await openai.chat.completions.create({
messages: [{ role: "system", content: "You are a helpful assistant." }],
model: "deepseek-v4-flash",
thinking: {"type": "enabled"},
reasoning_effort: "high",
stream: false,
});
console.log(completion.choices[0].message.content);
}
main();Control Thinking with reasoning_effort
Thinking mode is on by default. Two parameters control it: thinking toggles it, and reasoning_effort sets how much reasoning the model spends before answering.[6]
reasoning_effort supports three levels: low, high, and max. The default is high. For compatibility, low and medium map to high, and xhigh maps to max. Complex agent requests — like Claude Code or OpenCode — are automatically set to max.
When thinking is enabled, chain-of-thought reasoning comes back in the reasoning_content field. In multi-turn tool calls, you must return the assistant's reasoning_content together with the message, or the reasoning chain breaks.
- thinking: {"type": "enabled"} turns reasoning on; non-thinking mode turns it off.
- Default reasoning_effort is high.
- Reasoning output arrives in reasoning_content, matching the OpenAI o-series response shape.
| Requested value | Effective level |
|---|---|
| low / medium | high |
| high | high |
| max | max |
| xhigh | max |
Turn thinking off when you need the fastest possible response for simple tasks, and use max for hard coding and agent workloads.
Use the Responses API (Flash Only)
DeepSeek V4 Flash is the first DeepSeek model with native Responses API support — the same interface OpenAI's o-series uses.[7] Today only flash supports it; V4 Pro support is expected in early August 2026.
Streaming responses return semantic SSE events. Each event carries an event field and an increasing sequence_number, and the stream ends with response.completed, response.incomplete, or response.failed — there is no data: [DONE] message.
Unsupported parameters are silently ignored, so existing Responses API clients connect without code changes.
- Tools: function, plus server-side web_search and web_search_2025_08_26.
- custom tools are limited to {"type": "custom", "name": "apply_patch"} for Codex compatibility.
- Usage is reported in usage.input_tokens and usage.output_tokens, including cached and reasoning token breakdowns.
The base URL stays https://api.deepseek.com for the Responses API. Only the model and the endpoint differ from Chat Completions.
from openai import OpenAI
client = OpenAI(api_key="<your DeepSeek API Key>", base_url="https://api.deepseek.com")
response = client.responses.create(
model="deepseek-v4-flash",
instructions="You are a helpful assistant.",
input="Hi, how are you?",
)
print(response.output_text)Handle Errors, Rate Limits & Context Caching
DeepSeek documents seven HTTP error codes. Four mean fix your request and retry; three mean slow down and retry later.
Rate limits are measured by concurrency, not by requests per minute. Each account gets 2,500 concurrent requests for deepseek-v4-flash (500 for deepseek-v4-pro), and all API keys on the account share that pool. Exceeding it returns HTTP 429.[5]
Context caching is automatic and needs no configuration. Repeated prompt prefixes hit a server-side cache and are billed at the cache-hit price: $0.0028 per 1M tokens instead of $0.14 — roughly a 98% discount. Watch usage.prompt_cache_hit_tokens and usage.prompt_cache_miss_tokens in responses.
| Code | Meaning | What to do |
|---|---|---|
| 400 | Invalid Format | Fix the request body per the error message |
| 401 | Authentication Fails | Check or recreate your API key |
| 402 | Insufficient Balance | Top up on the billing page |
| 422 | Invalid Parameters | Fix the parameters per the error message |
| 429 | Rate Limit Reached | Slow your request pacing; consider a temporary switch |
| 500 | Server Error | Retry later; contact support if it persists |
| 503 | Server Overloaded | Retry later |
Caching is best-effort, so hit rates are not guaranteed. Keep stable system prompts and shared prefixes in front of your requests to maximize cache hits. If requests are slow or you hit 500/503, check the official API status page first — https://status.deepseek.com