Gemini 2.5 Flash Lite Preview API

google/gemini-2.5-flash-lite-preview
Gemini 2.5 Flash Lite Preview is a lightweight AI model developed by Google, optimized for quick responses and efficient processing, making it ideal for tasks requiring minimal latency and resource consumption
Context
1M tokens
Input
$0.0975 / 1M tokens
Output
$0.39 / 1M tokens
Released
Sep 9, 2025

How to use Gemini 2.5 Flash Lite Preview API

Install any OpenAI-compatible SDK, point it at api.aimlapi.com/v1, and set the model to google/gemini-2.5-flash-lite-preview.
import requests

r = requests.post(
    "https://api.aimlapi.com/v1/chat/completions",
    headers={"Authorization": "Bearer " + AIMLAPI_KEY},
    json={
      "model": "google/gemini-2.5-flash-lite-preview",
      "messages": [
        {
          "role": "user",
          "content": "Hello!"
        }
      ]
    },
)
print(r.json())
const r = await fetch("https://api.aimlapi.com/v1/chat/completions", {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.AIMLAPI_KEY}`,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
    "model": "google/gemini-2.5-flash-lite-preview",
    "messages": [
      {
        "role": "user",
        "content": "Hello!"
      }
    ]
  }),
});
console.log(await r.json());
curl -X POST https://api.aimlapi.com/v1/chat/completions \
  -H "Authorization: Bearer $AIMLAPI_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"google/gemini-2.5-flash-lite-preview","messages":[{"role":"user","content":"Hello!"}]}'

OpenAI-compatible — swap the base URL and it works with your existing SDK.

Gemini 2.5 Flash Lite Preview API Pricing

TypePrice
Input
$0.0975 / 1M tokens
Output
$0.39 / 1M tokens
Cached input
$0.0975 / 1M tokens

Preview models will typically have billing enabled and may come with more restrictive rate limits, and this exact model is documented with a 500,000,000 token rate limit.

Gemini 2.5 Flash Lite Preview Benchmarks

BenchmarkScoreWhat it measuresSourceRetrieved
Intelligence
10.4
Composite score across standardised reasoning, knowledge and problem-solving evaluations, measured independently by Artificial AnalysisSourceSeptember 12, 2026
Math
68.7
Composite score across standardised mathematics evaluations, measured independently by Artificial AnalysisSourceSeptember 12, 2026
Humanity's Last Exam
6.4%
Expert-level questions across many domainsSourceSeptember 21, 2026
GPQA Diamond
70.2%
Google-proof graduate science questions (hardest subset)SourceSeptember 21, 2026
AIME 2025
50.1%
Competition mathematics (AIME), 2025SourceSeptember 21, 2026

Gemini 2.5 Flash Lite Preview vs other models

ModelInputOutputContextBest for
$0.0975 / 1M tokens
$0.39 / 1M tokens
1M tokens
Agentic workflows and structured output
$0.325 / 1M tokens
$1.95 / 1M tokens
1M tokens
Reasoning + agents
$0.0975 / 1M tokens
$0.39 / 1M tokens
1M tokens
Reasoning + agents
$0.34385 / 1M tokens
$2.0631 / 1M tokens
1M tokens
Reasoning + agents

Frequently asked questions

Gemini 2.5 Flash Lite Preview has a 1,000,000 tokens context window and can return up to 1,048,576 tokens.

Gemini 2.5 Flash Lite Preview takes image, text as input and returns text.

Use google/gemini-2.5-flash-lite-preview as the model id. Requests go to https://api.aimlapi.com/v1/chat/completions.

Gemini 2.5 Flash Lite Preview became available on September 9, 2025.

Gemini 2.5 Flash Lite Preview is priced at input $0.0975 / 1M tokens, output $0.39 / 1M tokens, cached input $0.0975 / 1M tokens.

Yes, Gemini 2.5 Flash Lite Preview can stream responses as they are generated.

Yes, Gemini 2.5 Flash Lite Preview accepts image input alongside text.

Gemini 2.5 Flash Lite Preview was built by Google.

Yes, it supports tools including parallel tool calls and structured output.

Start building with Gemini 2.5 Flash Lite Preview

Get API Key
1000+ models, one API.