import requests, time headers = {"Authorization": "Bearer " + AIMLAPI_KEY} job = requests.post( "https://api.aimlapi.com/v2/video/generations", headers=headers, json={ "model": "x-ai/grok-imagine-video", "prompt": "A serene timelapse of clouds over a mountain range" }, ).json() gid = job["id"] while True: res = requests.get(f"https://api.aimlapi.com/v2/video/generations?generation_id={gid}", headers=headers).json() if res.get("status") in ("completed", "error"): break time.sleep(5) print(res["video"]["url"])
const headers = { Authorization: `Bearer ${process.env.AIMLAPI_KEY}`, "Content-Type": "application/json", }; const job = await (await fetch("https://api.aimlapi.com/v2/video/generations", { method: "POST", headers, body: JSON.stringify({ "model": "x-ai/grok-imagine-video", "prompt": "A serene timelapse of clouds over a mountain range" }), })).json(); let res; do { await new Promise((r) => setTimeout(r, 5000)); res = await (await fetch(`https://api.aimlapi.com/v2/video/generations?generation_id=${job.id}`, { headers })).json(); } while (!["completed", "error"].includes(res.status)); console.log(res.video.url);
# submit the generation — the response contains the job "id" curl -X POST https://api.aimlapi.com/v2/video/generations \ -H "Authorization: Bearer $AIMLAPI_KEY" \ -H "Content-Type: application/json" \ -d '{"model":"x-ai/grok-imagine-video","prompt":"A serene timelapse of clouds over a mountain range"}' # then poll until status is "completed" — video.url holds the result curl "https://api.aimlapi.com/v2/video/generations?generation_id={id}" -H "Authorization: Bearer $AIMLAPI_KEY"
OpenAI-compatible — swap the base URL and it works with your existing SDK.
| Type | Price |
|---|---|
| Output | |
| Benchmark | Score | What it measures | Source | Retrieved |
|---|---|---|---|---|
| Image-to-Video Arena | 1327 (Elo) | Human preference Elo from blind pairwise comparisons of videos animated from the same input image, measured independently by Artificial Analysis | Source | July 16, 2026 |
| Text-to-Video Arena | 1228 (Elo) | Human preference Elo from blind pairwise comparisons of generated videos, measured independently by Artificial Analysis | Source | July 16, 2026 |
| Model | Input | Output | Context | Best for |
|---|---|---|---|---|
Grok Imagine Video This page | Video generation | |||
| Video generation | ||||
| Video generation | ||||
| Video generation | ||||
| Cinematic video + native audio |
Grok Imagine Video takes image, text, video as input and returns video.
Grok Imagine Video is priced at $0.065 / sec (variable).
Grok Imagine Video is billed per generation — a fixed charge per output rather than by prompt length.
Yes, Grok Imagine Video accepts image input as well as a text prompt.
Grok Imagine Video was built by xAI .
Use x-ai/grok-imagine-video as the model id on AI/ML API.
Yes. Grok Imagine Video is served through AI/ML API, so the same key and endpoint format used for other models applies.
Yes, it takes image and text inputs together with video and produces video output.
It transforms static images into short animated video clips with coherent motion and visual consistency.
It is classified as a video generation model.