GLM 5.2 Fast Preview API

alibaba/glm-5.2-fast-preview
Latency-optimized GLM-5.2 preview — same 1M-token context, 1.5–2× faster output for real-time chat and agents.
Context
1M tokens
Input
$4.55 / 1M tokens
Output
$14.3 / 1M tokens
Released
Jul 10, 2026

How to use GLM 5.2 Fast Preview API

Install any OpenAI-compatible SDK, point it at api.aimlapi.com/v1, and set the model to alibaba/glm-5.2-fast-preview.
import requests

r = requests.post(
    "https://api.aimlapi.com/v1/chat/completions",
    headers={"Authorization": "Bearer " + AIMLAPI_KEY},
    json={
      "model": "alibaba/glm-5.2-fast-preview",
      "messages": [
        {
          "role": "user",
          "content": "Hello!"
        }
      ]
    },
)
print(r.json())
const r = await fetch("https://api.aimlapi.com/v1/chat/completions", {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.AIMLAPI_KEY}`,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
    "model": "alibaba/glm-5.2-fast-preview",
    "messages": [
      {
        "role": "user",
        "content": "Hello!"
      }
    ]
  }),
});
console.log(await r.json());
curl -X POST https://api.aimlapi.com/v1/chat/completions \
  -H "Authorization: Bearer $AIMLAPI_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"alibaba/glm-5.2-fast-preview","messages":[{"role":"user","content":"Hello!"}]}'

OpenAI-compatible — swap the base URL and it works with your existing SDK.

GLM 5.2 Fast Preview API Pricing

TypePrice
Input
$4.55 / 1M tokens
Output
$14.3 / 1M tokens
Cached input
$0.91 / 1M tokens

glm-5.2-fast-preview has no free quota on Alibaba Cloud’s Model Studio pricing page, and its input, output, and cached-input prices are listed separately for each deployment scope.

GLM 5.2 Fast Preview vs other models

ModelInputOutputContextBest for
$4.55 / 1M tokens
$14.3 / 1M tokens
1M tokens
Agentic workflows and structured output
$1.82 / 1M tokens
$5.72 / 1M tokens
1M tokens
Coding + agents
$1.82 / 1M tokens
$5.72 / 1M tokens
1M tokens
Reasoning + agents
$1.56 / 1M tokens
$5.2 / 1M tokens
262K tokens
Reasoning + agents

Frequently asked questions

GLM 5.2 Fast Preview has a 1,000,000 tokens context window and can return up to 131,072 tokens.

GLM 5.2 Fast Preview takes text as input and returns text.

Use alibaba/glm-5.2-fast-preview as the model id. Requests go to https://api.aimlapi.com/v1/chat/completions.

GLM 5.2 Fast Preview became available on July 10, 2026.

GLM 5.2 Fast Preview is priced at input $4.55 / 1M tokens, output $14.3 / 1M tokens, cached input $0.91 / 1M tokens.

Yes, GLM 5.2 Fast Preview can stream responses as they are generated.

GLM 5.2 Fast Preview was built by Zhipu AI.

It supports reasoning, streaming, structured output, tools, and parallel tool calls.

Start building with GLM 5.2 Fast Preview

Get API Key
1000+ models, one API.