import requests r = requests.post( "https://api.aimlapi.com/v1/chat/completions", headers={"Authorization": "Bearer " + AIMLAPI_KEY}, json={ "model": "thinkingmachines/inkling", "messages": [ { "role": "user", "content": "Hello!" } ] }, ) print(r.json())
const r = await fetch("https://api.aimlapi.com/v1/chat/completions", { method: "POST", headers: { Authorization: `Bearer ${process.env.AIMLAPI_KEY}`, "Content-Type": "application/json", }, body: JSON.stringify({ "model": "thinkingmachines/inkling", "messages": [ { "role": "user", "content": "Hello!" } ] }), }); console.log(await r.json());
curl -X POST https://api.aimlapi.com/v1/chat/completions \ -H "Authorization: Bearer $AIMLAPI_KEY" \ -H "Content-Type: application/json" \ -d '{"model":"thinkingmachines/inkling","messages":[{"role":"user","content":"Hello!"}]}'
OpenAI-compatible — swap the base URL and it works with your existing SDK.
| Type | Price |
|---|---|
| Input | |
| Output | |
| Cached input |
Inkling is offered at a 50% limited-time discount, and OpenRouter documentation also shows a free tier for thinkingmachines/inkling with 200 requests per day.
| Benchmark | Score | What it measures | Source | Retrieved |
|---|---|---|---|---|
| Intelligence | 25.5 | Composite score across standardised reasoning, knowledge and problem-solving evaluations, measured independently by Artificial Analysis | Source | September 12, 2026 |
| Coding | 52.1 | Composite score across standardised coding evaluations, measured independently by Artificial Analysis | Source | September 12, 2026 |
| GPQA Diamond | 87.2% | Google-proof graduate science questions (hardest subset) | Source | September 22, 2026 |
| SWE-bench Verified | 77.6% | Resolving verified real GitHub issues | Source | September 22, 2026 |
| MMAU | 77.2% | Multimodal audio understanding + reasoning | Source | September 22, 2026 |
| Model | Input | Output | Context | Best for |
|---|---|---|---|---|
Inkling This page | Agentic workflows and structured output | |||
| Efficient open-weight multimodal reasoning |
Inkling has a 1,048,576 tokens context window.
Inkling takes image, text as input and returns text.
Use thinkingmachines/inkling as the model id. Requests go to https://api.aimlapi.com/v1/chat/completions.
Inkling is priced at input $1.3754 / 1M tokens, output $5.57037 / 1M tokens, cached input $0.233818 / 1M tokens.
Yes, Inkling can stream responses as they are generated.
Yes, Inkling accepts image input alongside text.
Inkling was built by Thinking Machines.
Inkling is a mixture-of-experts model with 41B active parameters out of 975B total.
Inkling supports tools, parallel tool calls, structured output, and reasoning, making it suitable for agentic and tool-use systems.