import requests r = requests.post( "https://api.aimlapi.com/v1/chat/completions", headers={"Authorization": "Bearer " + AIMLAPI_KEY}, json={ "model": "deepseek/deepseek-v4.1-flash", "messages": [ { "role": "user", "content": "Hello!" } ] }, ) print(r.json())
const r = await fetch("https://api.aimlapi.com/v1/chat/completions", { method: "POST", headers: { Authorization: `Bearer ${process.env.AIMLAPI_KEY}`, "Content-Type": "application/json", }, body: JSON.stringify({ "model": "deepseek/deepseek-v4.1-flash", "messages": [ { "role": "user", "content": "Hello!" } ] }), }); console.log(await r.json());
curl -X POST https://api.aimlapi.com/v1/chat/completions \ -H "Authorization: Bearer $AIMLAPI_KEY" \ -H "Content-Type: application/json" \ -d '{"model":"deepseek/deepseek-v4.1-flash","messages":[{"role":"user","content":"Hello!"}]}'
OpenAI-compatible — swap the base URL and it works with your existing SDK.
| Type | Price |
|---|---|
| Input | |
| Output | |
| Cached input |
| Benchmark | Score | What it measures | Source | Retrieved |
|---|---|---|---|---|
| Intelligence | 39.5 | Composite score across standardised reasoning, knowledge and problem-solving evaluations, measured independently by Artificial Analysis | Source | September 12, 2026 |
| Model | Input | Output | Context | Best for |
|---|---|---|---|---|
DeepSeek V4.1 Flash This page | Reasoning + agents | |||
| Reasoning + agents | ||||
| Reasoning + agents | ||||
| Balanced coding + agents | ||||
| Long-context, multimodal & agentic workflows |
DeepSeek V4.1 Flash is the current build in DeepSeek's Flash line: a chat and reasoning model with a 1M-token context window, thinking mode enabled by default and image understanding.
DeepSeek reports that V4.1 Flash surpasses V4 Pro on performance, cost and speed. It also accepts image input, which V4 Pro does not.
Yes. It takes both image and text input and returns text, so screenshots, diagrams and scanned pages can go into the same request as your prompt.
Yes. Thinking mode is enabled by default, so the model reasons through a problem before answering unless you tell it otherwise.
1,048,576 tokens, with up to 384,000 tokens of output.
Yes, it supports function calling and structured outputs, alongside streaming responses.
Work that needs a very large window: whole-repository code analysis, long-document review, and agent loops that accumulate a lot of context. Image input widens that to mixed text-and-visual material.
Yes. The model also answers to the alias deepseek-flash.
Send a request to the chat completions endpoint with the model id deepseek/deepseek-v4.1-flash. The request shape is the same as for any other chat model on AI/ML API.