import requests r = requests.post( "https://api.aimlapi.com/v1/chat/completions", headers={"Authorization": "Bearer " + AIMLAPI_KEY}, json={ "model": "anthropic/claude-haiku-5.5", "messages": [ { "role": "user", "content": "Hello!" } ] }, ) print(r.json())
const r = await fetch("https://api.aimlapi.com/v1/chat/completions", { method: "POST", headers: { Authorization: "Bearer " + process.env.AIMLAPI_KEY, "Content-Type": "application/json", }, body: JSON.stringify({ "model": "anthropic/claude-haiku-5.5", "messages": [ { "role": "user", "content": "Hello!" } ] }), }); console.log(await r.json());
curl -X POST https://api.aimlapi.com/v1/chat/completions \ -H "Authorization: Bearer $AIMLAPI_KEY" \ -H "Content-Type: application/json" \ -d '{"model":"anthropic/claude-haiku-5.5","messages":[{"role":"user","content":"Hello!"}]}'
OpenAI-compatible — swap the base URL and it works with your existing SDK.
| Type | Price |
|---|---|
| Input | |
| Output | |
| Cached input |
Anthropic prices this model in two tiers by prompt length: the rates shown apply to prompts up to 100,000 tokens, and prompts above that are charged five times more ($0.65 in / $3.25 out per 1M tokens). Batch requests are half price.
| Benchmark | Score | What it measures | Source | Retrieved |
|---|---|---|---|---|
| GDPval-AA | 1620 | Economically valuable knowledge work across occupations, scored as an Elo-style rating | Source | October 7, 2026 |
| AA-Briefcase | 1578 | Multi-document office work: reading a brief, pulling the relevant facts and producing the deliverable | Source | October 7, 2026 |
| OSWorld | 72.4% | Computer-use across real desktop applications | Source | October 7, 2026 |
| Humanity's Last Exam | 45.9% | Expert-level questions across many domains | Source | October 7, 2026 |
| Terminal-Bench 4 | 39.2% | Autonomous shell and terminal task completion, fourth release | Source | October 7, 2026 |
| FrontierCode | 46.4% | Agentic coding on frontier-level software engineering tasks | Source | October 7, 2026 |
| Chartography | 46.4% | Visual reasoning over charts, plots and diagrams | Source | October 7, 2026 |
| Intelligence | 43.4 | Composite score across standardised reasoning, knowledge and problem-solving evaluations, measured independently by Artificial Analysis | Source | October 7, 2026 |
Figures published by Anthropic with the Claude Haiku 5.5 announcement (anthropic.com/claude-haiku-5-5), retrieved 7 October 2026. Vendor-reported and not independently reproduced. OSWorld 2.1 is the offline subset; Humanity's Last Exam is the no-tools score, rising to 57.4% with tools.
| Model | Input | Output | Context | Best for |
|---|---|---|---|---|
Claude Haiku 5.5 This page | High-volume, low-latency tasks | |||
| Multimodal reasoning | ||||
| Complex agentic coding and enterprise workflows | ||||
| Chat + assistants |
Claude Haiku 5.5 is Anthropic's fastest current model. Anthropic positions it for high-volume, latency-sensitive tasks such as classification, extraction and routing, and lists it as the fastest in the current lineup.
$0.13 per 1M input tokens and $0.65 per 1M output tokens through AI/ML API. Cached input reads are $0.013 per 1M tokens, following Anthropic's rate of 10% of the base input price.
1M tokens, with a maximum output of 128K tokens on the synchronous endpoint. Anthropic notes that batch requests can reach 300K output tokens with a beta header.
It uses adaptive thinking: the model decides how much to think, steered by the effort parameter. The default effort on Haiku 5.5 is medium.
Yes. Anthropic states that all current Claude models support text and image input, text output, multilingual capabilities, vision and tool use.
Both carry a 1M-token context window and 128K max output. Anthropic rates Haiku 5.5 as the fastest in the lineup and Sonnet 5.5 as the best combination of speed and intelligence. Sonnet 5.5 is listed at $2 / $10 per million tokens against Haiku's from $0.10 / from $0.50.