GLM 5.2 API

alibaba/glm-5.2
Zhipu AI's flagship MoE model — 1M-token context, built for agentic coding, tool use, and long-horizon reasoning.
Context
1M tokens
Input
$1.82 / 1M tokens
Output
$5.72 / 1M tokens
Released
Jun 16, 2026

How to use GLM 5.2 API

Install any OpenAI-compatible SDK, point it at api.aimlapi.com/v1, and set the model to alibaba/glm-5.2.
import requests

r = requests.post(
    "https://api.aimlapi.com/v1/chat/completions",
    headers={"Authorization": "Bearer " + AIMLAPI_KEY},
    json={
      "model": "alibaba/glm-5.2",
      "messages": [
        {
          "role": "user",
          "content": "Hello!"
        }
      ]
    },
)
print(r.json())
const r = await fetch("https://api.aimlapi.com/v1/chat/completions", {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.AIMLAPI_KEY}`,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
    "model": "alibaba/glm-5.2",
    "messages": [
      {
        "role": "user",
        "content": "Hello!"
      }
    ]
  }),
});
console.log(await r.json());
curl -X POST https://api.aimlapi.com/v1/chat/completions \
  -H "Authorization: Bearer $AIMLAPI_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"alibaba/glm-5.2","messages":[{"role":"user","content":"Hello!"}]}'

OpenAI-compatible — swap the base URL and it works with your existing SDK.

GLM 5.2 API Pricing

TypePrice
Input
$1.82 / 1M tokens
Output
$5.72 / 1M tokens
Cached input
$0.364 / 1M tokens

GLM-5.2’s documented pricing caveat is that cached input is billed at a discounted rate of $0.26 per million tokens, while Z.ai’s pricing also notes limited-time free cached-input storage.

GLM 5.2 Benchmarks

BenchmarkScoreWhat it measuresSourceRetrieved
Intelligence
34
Composite score across standardised reasoning, knowledge and problem-solving evaluations, measured independently by Artificial AnalysisSourceSeptember 12, 2026
Coding
68.8
Composite score across standardised coding evaluations, measured independently by Artificial AnalysisSourceSeptember 12, 2026
GPQA Diamond
91.2
Google-proof graduate science questions (hardest subset)SourceSeptember 21, 2026
Humanity's Last Exam
40.5
Expert-level questions across many domainsSourceSeptember 21, 2026

GLM 5.2 vs other models

ModelInputOutputContextBest for
GLM 5.2
This page
$1.82 / 1M tokens
$5.72 / 1M tokens
1M tokens
Agentic workflows and structured output
$1.82 / 1M tokens
$5.72 / 1M tokens
1M tokens
Coding + agents
$1.82 / 1M tokens
$5.72 / 1M tokens
1M tokens
Reasoning + agents
$1.56 / 1M tokens
$5.2 / 1M tokens
262K tokens
Reasoning + agents

Frequently asked questions

GLM 5.2 has a 1,000,000 tokens context window and can return up to 131,072 tokens.

GLM 5.2 takes text as input and returns text.

Use alibaba/glm-5.2 as the model id. Requests go to https://api.aimlapi.com/v1/chat/completions.

GLM 5.2 became available on June 16, 2026.

GLM 5.2 is priced at input $1.82 / 1M tokens, output $5.72 / 1M tokens, cached input $0.364 / 1M tokens.

Yes, GLM 5.2 can stream responses as they are generated.

GLM 5.2 was built by Zhipu AI.

Start building with GLM 5.2

Get API Key
1000+ models, one API.