Claude Haiku 5.5 API

anthropic/claude-haiku-5.5
Claude Haiku 5.5 is the fastest model in Anthropic's current lineup, positioned for high-volume, latency-sensitive tasks such as classification, extraction and routing. It carries a 1M-token context window and 128K maximum output.
Context
1M tokens
Input
$0.13 / 1M tokens
Output
$0.65 / 1M tokens
Released
Oct 7, 2026

How to use Claude Haiku 5.5 API

Install any OpenAI-compatible SDK, point it at api.aimlapi.com/v1, and set the model to anthropic/claude-haiku-5.5.
import requests

r = requests.post(
    "https://api.aimlapi.com/v1/chat/completions",
    headers={"Authorization": "Bearer " + AIMLAPI_KEY},
    json={
      "model": "anthropic/claude-haiku-5.5",
      "messages": [
        {
          "role": "user",
          "content": "Hello!"
        }
      ]
    },
)
print(r.json())
const r = await fetch("https://api.aimlapi.com/v1/chat/completions", {
  method: "POST",
  headers: {
    Authorization: "Bearer " + process.env.AIMLAPI_KEY,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
    "model": "anthropic/claude-haiku-5.5",
    "messages": [
      {
        "role": "user",
        "content": "Hello!"
      }
    ]
  }),
});
console.log(await r.json());
curl -X POST https://api.aimlapi.com/v1/chat/completions \
  -H "Authorization: Bearer $AIMLAPI_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"anthropic/claude-haiku-5.5","messages":[{"role":"user","content":"Hello!"}]}'

OpenAI-compatible — swap the base URL and it works with your existing SDK.

Claude Haiku 5.5 API Pricing

TypePrice
Input
$0.13 / 1M tokens
Output
$0.65 / 1M tokens
Cached input
$0.013 / 1M tokens

Anthropic prices this model in two tiers by prompt length: the rates shown apply to prompts up to 100,000 tokens, and prompts above that are charged five times more ($0.65 in / $3.25 out per 1M tokens). Batch requests are half price.

Claude Haiku 5.5 Benchmarks

BenchmarkScoreWhat it measuresSourceRetrieved
GDPval-AA
1620
Economically valuable knowledge work across occupations, scored as an Elo-style ratingSourceOctober 7, 2026
AA-Briefcase
1578
Multi-document office work: reading a brief, pulling the relevant facts and producing the deliverableSourceOctober 7, 2026
OSWorld
72.4%
Computer-use across real desktop applicationsSourceOctober 7, 2026
Humanity's Last Exam
45.9%
Expert-level questions across many domainsSourceOctober 7, 2026
Terminal-Bench 4
39.2%
Autonomous shell and terminal task completion, fourth releaseSourceOctober 7, 2026
FrontierCode
46.4%
Agentic coding on frontier-level software engineering tasksSourceOctober 7, 2026
Chartography
46.4%
Visual reasoning over charts, plots and diagramsSourceOctober 7, 2026
Intelligence
43.4
Composite score across standardised reasoning, knowledge and problem-solving evaluations, measured independently by Artificial AnalysisSourceOctober 7, 2026

Figures published by Anthropic with the Claude Haiku 5.5 announcement (anthropic.com/claude-haiku-5-5), retrieved 7 October 2026. Vendor-reported and not independently reproduced. OSWorld 2.1 is the offline subset; Humanity's Last Exam is the no-tools score, rising to 57.4% with tools.

Claude Haiku 5.5 vs other models

ModelInputOutputContextBest for
$0.13 / 1M tokens
$0.65 / 1M tokens
1M tokens
High-volume, low-latency tasks
$2.6 / 1M tokens
$13 / 1M tokens
1M tokens
Multimodal reasoning
$5.2 / 1M tokens
$26 / 1M tokens
1M tokens
Complex agentic coding and enterprise workflows
$6.877 / 1M tokens
$34.385 / 1M tokens
200K tokens
Chat + assistants

Frequently asked questions

Claude Haiku 5.5 is Anthropic's fastest current model. Anthropic positions it for high-volume, latency-sensitive tasks such as classification, extraction and routing, and lists it as the fastest in the current lineup.

$0.13 per 1M input tokens and $0.65 per 1M output tokens through AI/ML API. Cached input reads are $0.013 per 1M tokens, following Anthropic's rate of 10% of the base input price.

1M tokens, with a maximum output of 128K tokens on the synchronous endpoint. Anthropic notes that batch requests can reach 300K output tokens with a beta header.

It uses adaptive thinking: the model decides how much to think, steered by the effort parameter. The default effort on Haiku 5.5 is medium.

Yes. Anthropic states that all current Claude models support text and image input, text output, multilingual capabilities, vision and tool use.

Both carry a 1M-token context window and 128K max output. Anthropic rates Haiku 5.5 as the fastest in the lineup and Sonnet 5.5 as the best combination of speed and intelligence. Sonnet 5.5 is listed at $2 / $10 per million tokens against Haiku's from $0.10 / from $0.50.

Start building with Claude Haiku 5.5

Get API Key
1000+ models, one API.