Qwen3 VL 32B Thinking API

alibaba/qwen3-vl-32b-thinking
Qwen3 VL 32B Thinking is revolutionizing multimodal AI by enabling machines to process complex visual data alongside extended textual reasoning.
Context
126K tokens
Input
$0.91 / 1M tokens
Output
$10.92 / 1M tokens
Released
Nov 11, 2025

How to use Qwen3 VL 32B Thinking API

Install any OpenAI-compatible SDK, point it at api.aimlapi.com/v1, and set the model to alibaba/qwen3-vl-32b-thinking.
import requests

r = requests.post(
    "https://api.aimlapi.com/v1/chat/completions",
    headers={"Authorization": "Bearer " + AIMLAPI_KEY},
    json={
      "model": "alibaba/qwen3-vl-32b-thinking",
      "messages": [
        {
          "role": "user",
          "content": "Hello!"
        }
      ]
    },
)
print(r.json())
const r = await fetch("https://api.aimlapi.com/v1/chat/completions", {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.AIMLAPI_KEY}`,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
    "model": "alibaba/qwen3-vl-32b-thinking",
    "messages": [
      {
        "role": "user",
        "content": "Hello!"
      }
    ]
  }),
});
console.log(await r.json());
curl -X POST https://api.aimlapi.com/v1/chat/completions \
  -H "Authorization: Bearer $AIMLAPI_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"alibaba/qwen3-vl-32b-thinking","messages":[{"role":"user","content":"Hello!"}]}'

OpenAI-compatible — swap the base URL and it works with your existing SDK.

Qwen3 VL 32B Thinking API Pricing

TypePrice
Input
$0.91 / 1M tokens
Output
$10.92 / 1M tokens

Qwen3 VL 32B Thinking Benchmarks

BenchmarkScoreWhat it measuresSourceRetrieved
MMMU
78.1%
College-level multimodal understanding + reasoningSourceJuly 12, 2026
Intelligence
11.9
Composite score across standardised reasoning, knowledge and problem-solving evaluations, measured independently by Artificial AnalysisSourceSeptember 12, 2026
Math
84.7
Composite score across standardised mathematics evaluations, measured independently by Artificial AnalysisSourceSeptember 12, 2026

Qwen3 VL 32B Thinking vs other models

ModelInputOutputContextBest for
$0.91 / 1M tokens
$10.92 / 1M tokens
126K tokens
Coding + agents
$6.5 / 1M tokens
$39 / 1M tokens
1M tokens
Reasoning + agents
$2.6 / 1M tokens
$13 / 1M tokens
1M tokens
Balanced coding + agents
$3.9 / 1M tokens
$19.5 / 1M tokens
1M tokens
Long-context, multimodal & agentic workflows
$0.65 / 1M tokens
$3.9 / 1M tokens
1M tokens
Reasoning + agents

Frequently asked questions

Qwen3 VL 32B Thinking has a 126,000 tokens context window and can return up to 32,768 tokens.

Qwen3 VL 32B Thinking takes image, text as input and returns text.

Use alibaba/qwen3-vl-32b-thinking as the model id. Requests go to https://api.aimlapi.com/v1/chat/completions.

Qwen3 VL 32B Thinking became available on November 11, 2025.

Qwen3 VL 32B Thinking is priced at input $0.91 / 1M tokens, output $10.92 / 1M tokens.

Yes, Qwen3 VL 32B Thinking can stream responses as they are generated.

Yes, Qwen3 VL 32B Thinking accepts image input alongside text.

Qwen3 VL 32B Thinking was built by Alibaba Cloud.

Start building with Qwen3 VL 32B Thinking

Get API Key
1000+ models, one API.