import requests r = requests.post( "https://api.aimlapi.com/v1/images/generations", headers={"Authorization": "Bearer " + AIMLAPI_KEY}, json={ "model": "google/imagen4/preview", "prompt": "A cat holding a sign that says hello" }, ) print(r.json())
const r = await fetch("https://api.aimlapi.com/v1/images/generations", { method: "POST", headers: { Authorization: `Bearer ${process.env.AIMLAPI_KEY}`, "Content-Type": "application/json", }, body: JSON.stringify({ "model": "google/imagen4/preview", "prompt": "A cat holding a sign that says hello" }), }); console.log(await r.json());
curl -X POST https://api.aimlapi.com/v1/images/generations \ -H "Authorization: Bearer $AIMLAPI_KEY" \ -H "Content-Type: application/json" \ -d '{"model":"google/imagen4/preview","prompt":"A cat holding a sign that says hello"}'
OpenAI-compatible — swap the base URL and it works with your existing SDK.
| Type | Price |
|---|---|
| Output | |
| Benchmark | Score | What it measures | Source | Retrieved |
|---|---|---|---|---|
| Text-to-Image Arena | 1098 (Elo) | Human preference Elo from blind pairwise comparisons of generated images, measured independently by Artificial Analysis | Source | July 12, 2026 |
| Model | Input | Output | Context | Best for |
|---|---|---|---|---|
Imagen 4.0 Generate This page | Image generation at scale | |||
| Image generation at scale | ||||
| Image generation at scale | ||||
| Image generation at scale | ||||
| Image generation at scale |
Imagen 4.0 Generate takes image, text as input and returns image.
Imagen 4.0 Generate is priced at $0.052 / gen (variable).
Imagen 4.0 Generate is billed per generation — a fixed charge per output rather than by prompt length.
Yes, Imagen 4.0 Generate accepts image input as well as a text prompt.
Imagen 4.0 Generate was built by Google.
Send a request to https://api.aimlapi.com/v1/images/generations with google/imagen4/preview as the model id.
Yes. Imagen 4.0 Generate is served through AI/ML API, so the same key and endpoint format used for other models applies.
It is Google's text-to-image model focused on high-quality output with enhanced realism and understanding of prompts.