import requests r = requests.post( "https://api.aimlapi.com/v1/chat/completions", headers={"Authorization": "Bearer " + AIMLAPI_KEY}, json={ "model": "z-ai/glm-4.6v", "messages": [ { "role": "user", "content": "Hello!" } ] }, ) print(r.json())
const r = await fetch("https://api.aimlapi.com/v1/chat/completions", { method: "POST", headers: { Authorization: `Bearer ${process.env.AIMLAPI_KEY}`, "Content-Type": "application/json", }, body: JSON.stringify({ "model": "z-ai/glm-4.6v", "messages": [ { "role": "user", "content": "Hello!" } ] }), }); console.log(await r.json());
curl -X POST https://api.aimlapi.com/v1/chat/completions \ -H "Authorization: Bearer $AIMLAPI_KEY" \ -H "Content-Type: application/json" \ -d '{"model":"z-ai/glm-4.6v","messages":[{"role":"user","content":"Hello!"}]}'
OpenAI-compatible — swap the base URL and it works with your existing SDK.
| Type | Price |
|---|---|
| Input | |
| Output | |
| Cached input |
| Benchmark | Score | What it measures | Source | Retrieved |
|---|---|---|---|---|
| OSWorld | 37.2% | Computer-use across real desktop applications | Source | July 12, 2026 |
| MMMU | 76% | College-level multimodal understanding + reasoning | Source | July 12, 2026 |
| Intelligence | 11.2 | Composite score across standardised reasoning, knowledge and problem-solving evaluations, measured independently by Artificial Analysis | Source | September 12, 2026 |
| Math | 85.3 | Composite score across standardised mathematics evaluations, measured independently by Artificial Analysis | Source | September 12, 2026 |
| Model | Input | Output | Context | Best for |
|---|---|---|---|---|
GLM 4.6V This page | Reasoning + agents | |||
| Reasoning + agents | ||||
| Balanced coding + agents | ||||
| Long-context, multimodal & agentic workflows | ||||
| Reasoning + agents |
GLM 4.6V has a 131,072 tokens context window and can return up to 32,768 tokens.
GLM 4.6V takes image, text as input and returns text.
Use z-ai/glm-4.6v as the model id. Requests go to https://api.aimlapi.com/v1/chat/completions.
GLM 4.6V is priced at input $0.41262 / 1M tokens, output $1.23786 / 1M tokens, cached input $0.075647 / 1M tokens.
Yes, GLM 4.6V can stream responses as they are generated.
Yes, GLM 4.6V accepts image input alongside text.
GLM 4.6V was built by Zhipu AI.
Yes. It supports function calling, tool use, and parallel tool calls.