GLM 4.6V API

z-ai/glm-4.6v
GLM 4.6V on AIMLAPI.
Context
131K tokens
Input
$0.41262 / 1M tokens
Output
$1.23786 / 1M tokens

How to use GLM 4.6V API

Install any OpenAI-compatible SDK, point it at api.aimlapi.com/v1, and set the model to z-ai/glm-4.6v.
import requests

r = requests.post(
    "https://api.aimlapi.com/v1/chat/completions",
    headers={"Authorization": "Bearer " + AIMLAPI_KEY},
    json={
      "model": "z-ai/glm-4.6v",
      "messages": [
        {
          "role": "user",
          "content": "Hello!"
        }
      ]
    },
)
print(r.json())
const r = await fetch("https://api.aimlapi.com/v1/chat/completions", {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.AIMLAPI_KEY}`,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
    "model": "z-ai/glm-4.6v",
    "messages": [
      {
        "role": "user",
        "content": "Hello!"
      }
    ]
  }),
});
console.log(await r.json());
curl -X POST https://api.aimlapi.com/v1/chat/completions \
  -H "Authorization: Bearer $AIMLAPI_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"z-ai/glm-4.6v","messages":[{"role":"user","content":"Hello!"}]}'

OpenAI-compatible — swap the base URL and it works with your existing SDK.

GLM 4.6V API Pricing

TypePrice
Input
$0.41262 / 1M tokens
Output
$1.23786 / 1M tokens
Cached input
$0.075647 / 1M tokens

GLM 4.6V Benchmarks

BenchmarkScoreWhat it measuresSourceRetrieved
OSWorld
37.2%
Computer-use across real desktop applicationsSourceJuly 12, 2026
MMMU
76%
College-level multimodal understanding + reasoningSourceJuly 12, 2026
Intelligence
11.2
Composite score across standardised reasoning, knowledge and problem-solving evaluations, measured independently by Artificial AnalysisSourceSeptember 12, 2026
Math
85.3
Composite score across standardised mathematics evaluations, measured independently by Artificial AnalysisSourceSeptember 12, 2026

GLM 4.6V vs other models

ModelInputOutputContextBest for
GLM 4.6V
This page
$0.41262 / 1M tokens
$1.23786 / 1M tokens
131K tokens
Reasoning + agents
$6.5 / 1M tokens
$39 / 1M tokens
1.05M tokens
Reasoning + agents
$2.6 / 1M tokens
$13 / 1M tokens
1M tokens
Balanced coding + agents
$3.9 / 1M tokens
$19.5 / 1M tokens
1M tokens
Long-context, multimodal & agentic workflows
$0.65 / 1M tokens
$3.9 / 1M tokens
1.05M tokens
Reasoning + agents

Frequently asked questions

GLM 4.6V has a 131,072 tokens context window and can return up to 32,768 tokens.

GLM 4.6V takes image, text as input and returns text.

Use z-ai/glm-4.6v as the model id. Requests go to https://api.aimlapi.com/v1/chat/completions.

GLM 4.6V is priced at input $0.41262 / 1M tokens, output $1.23786 / 1M tokens, cached input $0.075647 / 1M tokens.

Yes, GLM 4.6V can stream responses as they are generated.

Yes, GLM 4.6V accepts image input alongside text.

GLM 4.6V was built by Zhipu AI.

Yes. It supports function calling, tool use, and parallel tool calls.

Start building with GLM 4.6V

Get API Key
1000+ models, one API.