Ling 3.1 Flash API

inclusionai/ling-3.1-flash
Ling 3.1 Flash is inclusionAI's hybrid-reasoning mixture-of-experts model: about 560B total parameters with roughly 25B active per token, served with a 256K context window.
Context
256K
Input
Output
Released
Sep 30, 2026

How to use Ling 3.1 Flash API

Install any OpenAI-compatible SDK, point it at api.aimlapi.com/v1, and set the model to inclusionai/ling-3.1-flash.
import requests

r = requests.post(
    "https://api.aimlapi.com/v1/chat/completions",
    headers={"Authorization": "Bearer " + AIMLAPI_KEY},
    json={
      "model": "inclusionai/ling-3.1-flash",
      "messages": [
        {
          "role": "user",
          "content": "Hello!"
        }
      ]
    },
)
print(r.json())
const r = await fetch("https://api.aimlapi.com/v1/chat/completions", {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.AIMLAPI_KEY}`,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
    "model": "inclusionai/ling-3.1-flash",
    "messages": [
      {
        "role": "user",
        "content": "Hello!"
      }
    ]
  }),
});
console.log(await r.json());
curl -X POST https://api.aimlapi.com/v1/chat/completions \
  -H "Authorization: Bearer $AIMLAPI_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"inclusionai/ling-3.1-flash","messages":[{"role":"user","content":"Hello!"}]}'

OpenAI-compatible — swap the base URL and it works with your existing SDK.

Ling 3.1 Flash API Pricing

TypePrice
Output

No per-token charge is listed for this model in the AI/ML API model feed at the moment.

Ling 3.1 Flash Benchmarks

BenchmarkScoreWhat it measuresSourceRetrieved
skillsBench
68.7%
Breadth of discrete professional skills carried out end to endSourceOctober 2, 2026
AutomationBench
52.5%
Completion of routine automation workflowsSourceOctober 2, 2026
CyberGym
87.9%
Security task completion in sandboxed environmentsSourceOctober 2, 2026
Finance Agent v2
57.9%
Financial research and analysis run as an agentSourceOctober 2, 2026
DRACO
85.5%
Data research and analysis with complex operationsSourceOctober 2, 2026
Terminal-Bench 4
40.4%
Autonomous shell and terminal task completion, fourth releaseSourceOctober 2, 2026
SWE-Atlas Codebase QnA
55.9%
Answering questions about an unfamiliar codebaseSourceOctober 2, 2026
HealthBench Professional
65.3%
Clinical and professional health question answeringSourceOctober 2, 2026

Figures reported by inclusionAI at launch; not independently reproduced.

Ling 3.1 Flash vs other models

ModelInputOutputContextBest for
256K
Low-latency assistance, extraction and routine automation with tool calling
$0.103155 / 1M tokens
$0.302588 / 1M tokens
256K tokens
Fast, low-cost reasoning and tool use at scale
$0.082524 / 1M tokens
$0.247572 / 1M tokens
128K tokens
Low-cost image and video understanding
$0.0975 / 1M
$0.8125 / 1M
262K tokens
Reasoning + agents

Frequently asked questions

Yes. The model was announced by inclusionAI on 30 September 2026 and is available through AI/ML API's chat completions endpoint as inclusionai/ling-3.1-flash. The model feed currently lists no per-token charge for it.

It is a mixture-of-experts model with roughly 560 billion total parameters, of which about 25 billion are active for any given token. That ratio is the point of the design: the model has the capacity of a very large network while only paying to run a small slice of it per token, which is what keeps latency low.

256K tokens as served at launch, with a maximum output of 32,768 tokens. inclusionAI has said it targets a context window of up to one million tokens; until that is actually served, 256K is the figure to plan against.

inclusionAI positions it for low-latency assistance, extraction and routine automation. Its published figures are mostly agentic rather than conversational — skill execution, automation, terminal work and codebase question answering, plus domain suites in security, finance and health — so it is aimed at pipelines that hand the model a toolset and let it work through a task.

No. Every figure on this page was reported by inclusionAI at launch and none has been reproduced by a third party. The model's weights are not published either, so independent evaluation is not yet possible. Treat the numbers as the vendor's own claims.

Ling 3.0 Flash is a far smaller mixture-of-experts model — roughly 124 billion total parameters with about 5 billion active — and it is available today. Ling 3.1 Flash raises both figures by roughly four to five times, to about 560 billion total and 25 billion active, and adds hybrid reasoning. Both are served with a 256K context window.

Start building with Ling 3.1 Flash

Get API Key
1000+ models, one API.