import requests r = requests.post( "https://api.aimlapi.com/v1/chat/completions", headers={"Authorization": "Bearer " + AIMLAPI_KEY}, json={ "model": "stepfun/step-5-preview", "messages": [ { "role": "user", "content": "Hello!" } ] }, ) print(r.json())
const r = await fetch("https://api.aimlapi.com/v1/chat/completions", { method: "POST", headers: { Authorization: `Bearer ${process.env.AIMLAPI_KEY}`, "Content-Type": "application/json", }, body: JSON.stringify({ "model": "stepfun/step-5-preview", "messages": [ { "role": "user", "content": "Hello!" } ] }), }); console.log(await r.json());
curl -X POST https://api.aimlapi.com/v1/chat/completions \ -H "Authorization: Bearer $AIMLAPI_KEY" \ -H "Content-Type: application/json" \ -d '{"model":"stepfun/step-5-preview","messages":[{"role":"user","content":"Hello!"}]}'
OpenAI-compatible — swap the base URL and it works with your existing SDK.
| Type | Price |
|---|---|
| Input | |
| Output | |
| Cached input |
Cache-miss input billing includes writing new content to the cache, output billing includes both reasoning tokens and the final answer, and API requests are rate-limited by account tier based on cumulative top-up amount.
| Benchmark | Score | What it measures | Source | Retrieved |
|---|---|---|---|---|
| Intelligence | 43.7 | Composite score across standardised reasoning, knowledge and problem-solving evaluations, measured independently by Artificial Analysis | Source | October 8, 2026 |
The primary published source is StepFun’s official Step 5 Preview announcement, “Step 5 Preview: Advancing the Pareto Frontier.” Source: https://www.stepfun.com/step-5-preview
| Model | Input | Output | Context | Best for |
|---|---|---|---|---|
Step 5 Preview This page | Agentic workflows and structured output | |||
| Coding + agents |
Step 5 Preview is StepFun’s flagship model for agentic work, with strengths in software engineering and professional knowledge work, particularly finance.
Step 5 Preview natively supports text, image, and video input.
Yes. Step 5 Preview can understand images and videos without requiring an additional vision model.
Yes. Step 5 Preview supports low, medium, and high reasoning effort levels.
Yes. Tool calling is supported for Step 5 Preview.
Yes. Step 5 Preview supports both JSON Mode and JSON Schema structured outputs.
Yes. Streaming output is supported for Step 5 Preview.
It can combine images, video, and text to extract and analyze multimodal information.
Yes. It is intended for tasks involving large amounts of information, tool calls, and continuous progress toward a deliverable.
Step 5 Preview uses a sparse Mixture-of-Experts architecture with 600 billion total parameters and 27 billion active parameters per token.