Mercury 2.5 API

inception/mercury-2.5
Mercury 2.5 is Inception's diffusion language model, generating and refining many tokens in parallel rather than one after another.
Context
256K tokens
Input
$0.27508 / 1M tokens
Output
$1.03155 / 1M tokens
Released
Sep 8, 2026

How to use Mercury 2.5 API

Install any OpenAI-compatible SDK, point it at api.aimlapi.com/v1, and set the model to inception/mercury-2.5.
import requests

r = requests.post(
    "https://api.aimlapi.com/v1/chat/completions",
    headers={"Authorization": "Bearer " + AIMLAPI_KEY},
    json={
      "model": "inception/mercury-2.5",
      "messages": [
        {
          "role": "user",
          "content": "Hello!"
        }
      ]
    },
)
print(r.json())
const r = await fetch("https://api.aimlapi.com/v1/chat/completions", {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.AIMLAPI_KEY}`,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
    "model": "inception/mercury-2.5",
    "messages": [
      {
        "role": "user",
        "content": "Hello!"
      }
    ]
  }),
});
console.log(await r.json());
curl -X POST https://api.aimlapi.com/v1/chat/completions \
  -H "Authorization: Bearer $AIMLAPI_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"inception/mercury-2.5","messages":[{"role":"user","content":"Hello!"}]}'

OpenAI-compatible — swap the base URL and it works with your existing SDK.

Mercury 2.5 API Pricing

TypePrice
Input
$0.27508 / 1M tokens
Output
$1.03155 / 1M tokens
Cached input
$0.027508 / 1M tokens

Mercury 2.5 vs other models

ModelInputOutputContextBest for
Mercury 2.5
This page
$0.27508 / 1M tokens
$1.03155 / 1M tokens
256K tokens
Low-latency reasoning and agent loops
$0.34385 / 1M tokens
$1.03155 / 1M tokens
128K tokens
Reasoning + agents
$6.5 / 1M tokens
$39 / 1M tokens
1M tokens
Reasoning + agents
$2.6 / 1M tokens
$13 / 1M tokens
1M tokens
Balanced coding + agents
$1.95 / 1M tokens
$11.7 / 1M tokens
1M tokens
Reasoning + agents

Frequently asked questions

Mercury 2.5 is a diffusion language model from Inception. Unlike models that emit one token after another, it generates and refines many tokens in parallel.

A standard LLM predicts tokens one at a time, each conditioned on the previous one. A diffusion language model starts from a rough draft of the whole output and refines it over several passes, so many tokens improve at once.

Inception reports throughput of 1,107 tokens per second on widely available NVIDIA GPUs, and describes it as the fastest reasoning model in production.

Yes. It supports function calling and can issue parallel tool calls within a single turn.

Yes. It supports schema-aligned JSON output, so responses can be constrained to the shape your code expects.

Yes. Reasoning depth is tunable, so you can trade latency against deliberation on a per-request basis.

Mercury 2.5 takes text and image input and returns text. It also supports file input and web search.

Send a request to the chat completions endpoint with the model id inception/mercury-2.5. The request shape is the same as for any other chat model on AI/ML API.

Inception reports roughly a 40% increase in intelligence over Mercury 2, while keeping the parallel decoding that makes the family fast.

Latency-sensitive work such as agent loops that make many sequential calls, and high-volume generation where throughput is the constraint.

Start building with Mercury 2.5

Get API Key
1000+ models, one API.