import requests r = requests.post( "https://api.aimlapi.com/v1/chat/completions", headers={"Authorization": "Bearer " + AIMLAPI_KEY}, json={ "model": "google/gemini-2.5-flash-lite-preview", "messages": [ { "role": "user", "content": "Hello!" } ] }, ) print(r.json())
const r = await fetch("https://api.aimlapi.com/v1/chat/completions", { method: "POST", headers: { Authorization: `Bearer ${process.env.AIMLAPI_KEY}`, "Content-Type": "application/json", }, body: JSON.stringify({ "model": "google/gemini-2.5-flash-lite-preview", "messages": [ { "role": "user", "content": "Hello!" } ] }), }); console.log(await r.json());
curl -X POST https://api.aimlapi.com/v1/chat/completions \ -H "Authorization: Bearer $AIMLAPI_KEY" \ -H "Content-Type: application/json" \ -d '{"model":"google/gemini-2.5-flash-lite-preview","messages":[{"role":"user","content":"Hello!"}]}'
OpenAI-compatible — swap the base URL and it works with your existing SDK.
| Type | Price |
|---|---|
| Input | |
| Output | |
| Cached input |
Preview models will typically have billing enabled and may come with more restrictive rate limits, and this exact model is documented with a 500,000,000 token rate limit.
| Benchmark | Score | What it measures | Source | Retrieved |
|---|---|---|---|---|
| Intelligence | 10.4 | Composite score across standardised reasoning, knowledge and problem-solving evaluations, measured independently by Artificial Analysis | Source | September 12, 2026 |
| Math | 68.7 | Composite score across standardised mathematics evaluations, measured independently by Artificial Analysis | Source | September 12, 2026 |
| Humanity's Last Exam | 6.4% | Expert-level questions across many domains | Source | September 21, 2026 |
| GPQA Diamond | 70.2% | Google-proof graduate science questions (hardest subset) | Source | September 21, 2026 |
| AIME 2025 | 50.1% | Competition mathematics (AIME), 2025 | Source | September 21, 2026 |
| Model | Input | Output | Context | Best for |
|---|---|---|---|---|
Gemini 2.5 Flash Lite Preview This page | Agentic workflows and structured output | |||
| Reasoning + agents | ||||
| Reasoning + agents | ||||
| Reasoning + agents |
Gemini 2.5 Flash Lite Preview has a 1,000,000 tokens context window and can return up to 1,048,576 tokens.
Gemini 2.5 Flash Lite Preview takes image, text as input and returns text.
Use google/gemini-2.5-flash-lite-preview as the model id. Requests go to https://api.aimlapi.com/v1/chat/completions.
Gemini 2.5 Flash Lite Preview became available on September 9, 2025.
Gemini 2.5 Flash Lite Preview is priced at input $0.0975 / 1M tokens, output $0.39 / 1M tokens, cached input $0.0975 / 1M tokens.
Yes, Gemini 2.5 Flash Lite Preview can stream responses as they are generated.
Yes, Gemini 2.5 Flash Lite Preview accepts image input alongside text.
Gemini 2.5 Flash Lite Preview was built by Google.
Yes, it supports tools including parallel tool calls and structured output.