Granite 4.1 8B API

ibm-granite/granite-4.1-8b
Granite 4.1 8B on AIMLAPI.
Context
131K tokens
Input
$0.06877 / 1M tokens
Output
$0.13754 / 1M tokens

How to use Granite 4.1 8B API

Install any OpenAI-compatible SDK, point it at api.aimlapi.com/v1, and set the model to ibm-granite/granite-4.1-8b.
import requests

r = requests.post(
    "https://api.aimlapi.com/v1/chat/completions",
    headers={"Authorization": "Bearer " + AIMLAPI_KEY},
    json={
      "model": "ibm-granite/granite-4.1-8b",
      "messages": [
        {
          "role": "user",
          "content": "Hello!"
        }
      ]
    },
)
print(r.json())
const r = await fetch("https://api.aimlapi.com/v1/chat/completions", {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.AIMLAPI_KEY}`,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
    "model": "ibm-granite/granite-4.1-8b",
    "messages": [
      {
        "role": "user",
        "content": "Hello!"
      }
    ]
  }),
});
console.log(await r.json());
curl -X POST https://api.aimlapi.com/v1/chat/completions \
  -H "Authorization: Bearer $AIMLAPI_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"ibm-granite/granite-4.1-8b","messages":[{"role":"user","content":"Hello!"}]}'

OpenAI-compatible — swap the base URL and it works with your existing SDK.

Granite 4.1 8B API Pricing

TypePrice
Input
$0.06877 / 1M tokens
Output
$0.13754 / 1M tokens
Cached input
$0.06877 / 1M tokens

Granite 4.1 8B Benchmarks

BenchmarkScoreWhat it measuresSourceRetrieved
Intelligence
6.6
Composite score across standardised reasoning, knowledge and problem-solving evaluations, measured independently by Artificial AnalysisSourceSeptember 12, 2026
Coding
9.5
Composite score across standardised coding evaluations, measured independently by Artificial AnalysisSourceSeptember 12, 2026

Granite 4.1 8B vs other models

ModelInputOutputContextBest for
$0.06877 / 1M tokens
$0.13754 / 1M tokens
131K tokens
Reasoning + agents
$6.5 / 1M tokens
$39 / 1M tokens
1.05M tokens
Reasoning + agents
$2.6 / 1M tokens
$13 / 1M tokens
1M tokens
Balanced coding + agents
$3.9 / 1M tokens
$19.5 / 1M tokens
1M tokens
Long-context, multimodal & agentic workflows
$0.65 / 1M tokens
$3.9 / 1M tokens
1.05M tokens
Reasoning + agents

Frequently asked questions

Granite 4.1 8B has a 131,072 tokens context window and can return up to 131,072 tokens.

Granite 4.1 8B takes image, text as input and returns text.

Use ibm-granite/granite-4.1-8b as the model id. Requests go to https://api.aimlapi.com/v1/chat/completions.

Granite 4.1 8B is priced at input $0.06877 / 1M tokens, output $0.13754 / 1M tokens, cached input $0.06877 / 1M tokens.

Yes, Granite 4.1 8B can stream responses as they are generated.

Yes, Granite 4.1 8B accepts image input alongside text.

It is a chat model supporting reasoning, structured outputs, file input, and web search, suited for conversational and tool-assisted tasks.

Yes, it supports function calling, tool use, and parallel tool calls.

Start building with Granite 4.1 8B

Get API Key
1000+ models, one API.