GLM-5.3-Flash API

z-ai/glm-5.3-flash
Zhipu's cost-efficient Flash variant of GLM-5.3 -- the model previously served anonymously as the stealth codename "Ox Alpha."
Context
1M tokens
Input
$0.195 / 1M tokens
Output
$0.325 / 1M tokens
Released
Aug 26, 2026

How to use GLM-5.3-Flash API

Install any OpenAI-compatible SDK, point it at api.aimlapi.com/v1, and set the model to z-ai/glm-5.3-flash.
import requests

r = requests.post(
    "https://api.aimlapi.com/v1/chat/completions",
    headers={"Authorization": "Bearer " + AIMLAPI_KEY},
    json={
      "model": "z-ai/glm-5.3-flash",
      "messages": [
        {
          "role": "user",
          "content": "Hello!"
        }
      ]
    },
)
print(r.json())
const r = await fetch("https://api.aimlapi.com/v1/chat/completions", {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.AIMLAPI_KEY}`,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
    "model": "z-ai/glm-5.3-flash",
    "messages": [
      {
        "role": "user",
        "content": "Hello!"
      }
    ]
  }),
});
console.log(await r.json());
curl -X POST https://api.aimlapi.com/v1/chat/completions \
-H "Authorization: Bearer $AIMLAPI_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"z-ai/glm-5.3-flash","messages":[{"role":"user","content":"Hello!"}]}'

OpenAI-compatible — swap the base URL and it works with your existing SDK.

GLM-5.3-Flash API Pricing

TypePrice
Input
$0.195 / 1M tokens
Output
$0.325 / 1M tokens
Cached input
$0.039 / 1M tokens

MIT-licensed open weights also available on Hugging Face.

GLM-5.3-Flash Benchmarks

BenchmarkScoreWhat it measuresSourceRetrieved
Intelligence
41.9
Composite score across standardised reasoning, knowledge and problem-solving evaluations, measured independently by Artificial AnalysisSourceSeptember 12, 2026
Coding
71.5
Composite score across standardised coding evaluations, measured independently by Artificial AnalysisSourceSeptember 12, 2026

GLM-5.3-Flash vs other models

ModelInputOutputContextBest for
GLM-5.3-Flash
This page
$0.195 / 1M tokens
$0.325 / 1M tokens
1M tokens
Reasoning + agents
$6.5 / 1M tokens
$39 / 1M tokens
1.05M tokens
Reasoning + agents
$2.6 / 1M tokens
$13 / 1M tokens
1M tokens
Balanced coding + agents
$3.9 / 1M tokens
$19.5 / 1M tokens
1M tokens
Long-context, multimodal & agentic workflows
$0.65 / 1M tokens
$3.9 / 1M tokens
1.05M tokens
Reasoning + agents

Frequently asked questions

GLM-5.3-Flash has a 1,048,576 tokens context window and can return up to 131,072 tokens.

GLM-5.3-Flash takes image, text, video as input and returns text.

Use z-ai/glm-5.3-flash as the model id. Requests go to https://api.aimlapi.com/v1/chat/completions.

GLM-5.3-Flash became available on August 26, 2026.

GLM-5.3-Flash is priced at input $0.195 / 1M tokens, output $0.325 / 1M tokens, cached input $0.039 / 1M tokens.

Yes, GLM-5.3-Flash can stream responses as they are generated.

Yes, GLM-5.3-Flash accepts image input alongside text.

GLM-5.3-Flash was built by Zhipu AI.

Start building with GLM-5.3-Flash

Get API Key
1000+ models, one API.