import requests r = requests.post( "https://api.aimlapi.com/v1/chat/completions", headers={"Authorization": "Bearer " + AIMLAPI_KEY}, json={ "model": "google/gemini-3-1-flash-lite", "messages": [ { "role": "user", "content": "Hello!" } ] }, ) print(r.json())
const r = await fetch("https://api.aimlapi.com/v1/chat/completions", { method: "POST", headers: { Authorization: `Bearer ${process.env.AIMLAPI_KEY}`, "Content-Type": "application/json", }, body: JSON.stringify({ "model": "google/gemini-3-1-flash-lite", "messages": [ { "role": "user", "content": "Hello!" } ] }), }); console.log(await r.json());
curl -X POST https://api.aimlapi.com/v1/chat/completions \ -H "Authorization: Bearer $AIMLAPI_KEY" \ -H "Content-Type: application/json" \ -d '{"model":"google/gemini-3-1-flash-lite","messages":[{"role":"user","content":"Hello!"}]}'
OpenAI-compatible — swap the base URL and it works with your existing SDK.
| Type | Price |
|---|---|
| Input | |
| Output | |
| Cached input |
| Benchmark | Score | What it measures | Source | Retrieved |
|---|---|---|---|---|
| LiveCodeBench | 72% | Contamination-free competitive programming problems | Source | July 12, 2026 |
| AIME 2025 | 16.7% | Competition mathematics (AIME), 2025 | Source | July 12, 2026 |
| GPQA Diamond | 86.9% | Google-proof graduate science questions (hardest subset) | Source | July 12, 2026 |
| Intelligence | 16 | Composite score across standardised reasoning, knowledge and problem-solving evaluations, measured independently by Artificial Analysis | Source | September 12, 2026 |
| Coding | 34.7 | Composite score across standardised coding evaluations, measured independently by Artificial Analysis | Source | September 12, 2026 |
| Model | Input | Output | Context | Best for |
|---|---|---|---|---|
Gemini 3.1 Flash Lite This page | Reasoning + agents | |||
| Reasoning + agents | ||||
| Balanced coding + agents | ||||
| Long-context, multimodal & agentic workflows | ||||
| Reasoning + agents |
Gemini 3.1 Flash Lite has a 1,000,000 tokens context window and can return up to 65,536 tokens.
Gemini 3.1 Flash Lite takes audio, image, text, video as input and returns audio, text.
Use google/gemini-3-1-flash-lite as the model id. Requests go to https://api.aimlapi.com/v1/chat/completions.
Gemini 3.1 Flash Lite is priced at input $0.325 / 1M tokens, output $1.95 / 1M tokens, cached input $0.0325 / 1M tokens.
Yes, Gemini 3.1 Flash Lite can stream responses as they are generated.
Yes, Gemini 3.1 Flash Lite accepts image input alongside text.
Gemini 3.1 Flash Lite was built by Google.
Yes, it includes reasoning as one of its capabilities.
It is designed for cost-efficient, high-volume tasks such as translation, lightweight reasoning, and simple agent workflows.
Yes, it supports function calling, tools, and parallel tool calls.