DeepSeek V4 Flash API

deepseek/deepseek-v4-flash
DeepSeek V4 Flash and Pro are large language models for chat and reasoning tasks with up to 1M context length.
Context
1M tokens
Input
$0.39 / 1M tokens
Output
$1.56 / 1M tokens
Released
Apr 24, 2026

How to use DeepSeek V4 Flash API

Install any OpenAI-compatible SDK, point it at api.aimlapi.com/v1, and set the model to deepseek/deepseek-v4-flash.
import requests

r = requests.post(
    "https://api.aimlapi.com/v1/chat/completions",
    headers={"Authorization": "Bearer " + AIMLAPI_KEY},
    json={
      "model": "deepseek/deepseek-v4-flash",
      "messages": [
        {
          "role": "user",
          "content": "Hello!"
        }
      ]
    },
)
print(r.json())
const r = await fetch("https://api.aimlapi.com/v1/chat/completions", {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.AIMLAPI_KEY}`,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
    "model": "deepseek/deepseek-v4-flash",
    "messages": [
      {
        "role": "user",
        "content": "Hello!"
      }
    ]
  }),
});
console.log(await r.json());
curl -X POST https://api.aimlapi.com/v1/chat/completions \
  -H "Authorization: Bearer $AIMLAPI_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"deepseek/deepseek-v4-flash","messages":[{"role":"user","content":"Hello!"}]}'

OpenAI-compatible — swap the base URL and it works with your existing SDK.

DeepSeek V4 Flash API Pricing

TypePrice
Input
$0.39 / 1M tokens
Output
$1.56 / 1M tokens
Cached input
$0.0078 / 1M tokens

DeepSeek V4 Flash Benchmarks

BenchmarkScoreWhat it measuresSourceRetrieved
Terminal-Bench
56.9%
Autonomous shell/terminal task completionSourceJuly 12, 2026
LiveCodeBench
91.6%
Contamination-free competitive programming problemsSourceJuly 12, 2026
GPQA Diamond
88.1%
Google-proof graduate science questions (hardest subset)SourceJuly 12, 2026
SWE-bench Verified
79%
Resolving verified real GitHub issuesSourceJuly 12, 2026
Intelligence
34.5
Composite score across standardised reasoning, knowledge and problem-solving evaluations, measured independently by Artificial AnalysisSourceSeptember 12, 2026
Coding
69.1
Composite score across standardised coding evaluations, measured independently by Artificial AnalysisSourceSeptember 12, 2026

DeepSeek V4 Flash vs other models

ModelInputOutputContextBest for
$0.39 / 1M tokens
$1.56 / 1M tokens
1M tokens
Reasoning + agents
$6.5 / 1M tokens
$39 / 1M tokens
1.05M tokens
Reasoning + agents
$2.6 / 1M tokens
$13 / 1M tokens
1M tokens
Balanced coding + agents
$3.9 / 1M tokens
$19.5 / 1M tokens
1M tokens
Long-context, multimodal & agentic workflows
$0.65 / 1M tokens
$3.9 / 1M tokens
1.05M tokens
Reasoning + agents

Frequently asked questions

DeepSeek V4 Flash has a 1,000,000 tokens context window and can return up to 384,000 tokens.

DeepSeek V4 Flash takes text as input and returns text.

Use deepseek/deepseek-v4-flash as the model id. Requests go to https://api.aimlapi.com/v1/chat/completions.

DeepSeek V4 Flash became available on April 24, 2026.

DeepSeek V4 Flash is priced at input $0.39 / 1M tokens, output $1.56 / 1M tokens, cached input $0.0078 / 1M tokens.

Yes, DeepSeek V4 Flash can stream responses as they are generated.

DeepSeek V4 Flash was built by DeepSeek.

It is a chat and reasoning model supporting both standard and deeper thinking modes for various language tasks.

Start building with DeepSeek V4 Flash

Get API Key
1000+ models, one API.