Inkling API

thinkingmachines/inkling
Thinking Machines' open-weight multimodal MoE model — 41B active of 975B total parameters, built for reasoning, coding, and agentic tool use.
Context
1.0M tokens
Input
$1.3754 / 1M tokens
Output
$5.57037 / 1M tokens

How to use Inkling API

Install any OpenAI-compatible SDK, point it at api.aimlapi.com/v1, and set the model to thinkingmachines/inkling.
import requests

r = requests.post(
    "https://api.aimlapi.com/v1/chat/completions",
    headers={"Authorization": "Bearer " + AIMLAPI_KEY},
    json={
      "model": "thinkingmachines/inkling",
      "messages": [
        {
          "role": "user",
          "content": "Hello!"
        }
      ]
    },
)
print(r.json())
const r = await fetch("https://api.aimlapi.com/v1/chat/completions", {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.AIMLAPI_KEY}`,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
    "model": "thinkingmachines/inkling",
    "messages": [
      {
        "role": "user",
        "content": "Hello!"
      }
    ]
  }),
});
console.log(await r.json());
curl -X POST https://api.aimlapi.com/v1/chat/completions \
  -H "Authorization: Bearer $AIMLAPI_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"thinkingmachines/inkling","messages":[{"role":"user","content":"Hello!"}]}'

OpenAI-compatible — swap the base URL and it works with your existing SDK.

Inkling API Pricing

TypePrice
Input
$1.3754 / 1M tokens
Output
$5.57037 / 1M tokens
Cached input
$0.233818 / 1M tokens

Inkling is offered at a 50% limited-time discount, and OpenRouter documentation also shows a free tier for thinkingmachines/inkling with 200 requests per day.

Inkling Benchmarks

BenchmarkScoreWhat it measuresSourceRetrieved
Intelligence
25.5
Composite score across standardised reasoning, knowledge and problem-solving evaluations, measured independently by Artificial AnalysisSourceSeptember 12, 2026
Coding
52.1
Composite score across standardised coding evaluations, measured independently by Artificial AnalysisSourceSeptember 12, 2026
GPQA Diamond
87.2%
Google-proof graduate science questions (hardest subset)SourceSeptember 22, 2026
SWE-bench Verified
77.6%
Resolving verified real GitHub issuesSourceSeptember 22, 2026
MMAU
77.2%
Multimodal audio understanding + reasoningSourceSeptember 22, 2026

Inkling vs other models

ModelInputOutputContextBest for
Inkling
This page
$1.3754 / 1M tokens
$5.57037 / 1M tokens
1.0M tokens
Agentic workflows and structured output
$0.797732 / 1M tokens
$1.980576 / 1M tokens
524K tokens
Efficient open-weight multimodal reasoning

Frequently asked questions

Inkling has a 1,048,576 tokens context window.

Inkling takes image, text as input and returns text.

Use thinkingmachines/inkling as the model id. Requests go to https://api.aimlapi.com/v1/chat/completions.

Inkling is priced at input $1.3754 / 1M tokens, output $5.57037 / 1M tokens, cached input $0.233818 / 1M tokens.

Yes, Inkling can stream responses as they are generated.

Yes, Inkling accepts image input alongside text.

Inkling was built by Thinking Machines.

Inkling is a mixture-of-experts model with 41B active parameters out of 975B total.

Inkling supports tools, parallel tool calls, structured output, and reasoning, making it suitable for agentic and tool-use systems.

Start building with Inkling

Get API Key
1000+ models, one API.