import requests r = requests.post( "https://api.aimlapi.com/v1/chat/completions", headers={"Authorization": "Bearer " + AIMLAPI_KEY}, json={ "model": "alibaba/qwen3-235b-a22b-thinking-2507", "messages": [ { "role": "user", "content": "Hello!" } ] }, ) print(r.json())
const r = await fetch("https://api.aimlapi.com/v1/chat/completions", { method: "POST", headers: { Authorization: `Bearer ${process.env.AIMLAPI_KEY}`, "Content-Type": "application/json", }, body: JSON.stringify({ "model": "alibaba/qwen3-235b-a22b-thinking-2507", "messages": [ { "role": "user", "content": "Hello!" } ] }), }); console.log(await r.json());
curl -X POST https://api.aimlapi.com/v1/chat/completions \ -H "Authorization: Bearer $AIMLAPI_KEY" \ -H "Content-Type: application/json" \ -d '{"model":"alibaba/qwen3-235b-a22b-thinking-2507","messages":[{"role":"user","content":"Hello!"}]}'
OpenAI-compatible — swap the base URL and it works with your existing SDK.
| Type | Price |
|---|---|
| Input | |
| Output | |
| Benchmark | Score | What it measures | Source | Retrieved |
|---|---|---|---|---|
| LiveCodeBench | 74.1% | Contamination-free competitive programming problems | Source | July 8, 2026 |
| GPQA Diamond | 81.1% | Google-proof graduate science questions (hardest subset) | Source | July 8, 2026 |
| AIME 2025 | 92.3% | Competition mathematics (AIME), 2025 | Source | July 8, 2026 |
| Intelligence | 12.7 | Composite score across standardised reasoning, knowledge and problem-solving evaluations, measured independently by Artificial Analysis | Source | September 12, 2026 |
| Coding | 22.1 | Composite score across standardised coding evaluations, measured independently by Artificial Analysis | Source | September 12, 2026 |
| Math | 91 | Composite score across standardised mathematics evaluations, measured independently by Artificial Analysis | Source | September 12, 2026 |
| Model | Input | Output | Context | Best for |
|---|---|---|---|---|
Qwen3 Thinking 2507 This page | Coding + agents | |||
| Reasoning + agents | ||||
| Balanced coding + agents | ||||
| Long-context, multimodal & agentic workflows | ||||
| Reasoning + agents |
Qwen3 Thinking 2507 has a 32,000 tokens context window and can return up to 16,000 tokens.
Qwen3 Thinking 2507 takes text as input and returns text.
Use alibaba/qwen3-235b-a22b-thinking-2507 as the model id. Requests go to https://api.aimlapi.com/v1/chat/completions.
Qwen3 Thinking 2507 became available on September 9, 2025.
Qwen3 Thinking 2507 is priced at input $0.299 / 1M tokens, output $2.99 / 1M tokens.
Yes, Qwen3 Thinking 2507 can stream responses as they are generated.
Qwen3 Thinking 2507 was built by Alibaba Cloud.
Yes, it supports streaming responses along with tool use, including parallel tool calls.