Gemini 3.5 Flash Lite API

google/gemini-3.5-flash-lite
Google's most cost-efficient GA model for high-volume agentic and multimodal tasks.
Context
1.05M tokens
Input
$0.39 / 1M tokens
Output
$3.25 / 1M tokens
Released
Jul 21, 2026

How to use Gemini 3.5 Flash Lite API

Install any OpenAI-compatible SDK, point it at api.aimlapi.com/v1, and set the model to google/gemini-3.5-flash-lite.
import requests

r = requests.post(
    "https://api.aimlapi.com/v1/chat/completions",
    headers={"Authorization": "Bearer " + AIMLAPI_KEY},
    json={
      "model": "google/gemini-3.5-flash-lite",
      "messages": [
        {
          "role": "user",
          "content": "Hello!"
        }
      ]
    },
)
print(r.json())
const r = await fetch("https://api.aimlapi.com/v1/chat/completions", {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.AIMLAPI_KEY}`,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
    "model": "google/gemini-3.5-flash-lite",
    "messages": [
      {
        "role": "user",
        "content": "Hello!"
      }
    ]
  }),
});
console.log(await r.json());
curl -X POST https://api.aimlapi.com/v1/chat/completions \
  -H "Authorization: Bearer $AIMLAPI_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"google/gemini-3.5-flash-lite","messages":[{"role":"user","content":"Hello!"}]}'

OpenAI-compatible — swap the base URL and it works with your existing SDK.

Gemini 3.5 Flash Lite API Pricing

TypePrice
Input
$0.39 / 1M tokens
Output
$3.25 / 1M tokens
Cached input
$0.039 / 1M tokens

Google Search grounding billed separately at $0.0455 per call. Cached input priced at $0.039/1M.

Gemini 3.5 Flash Lite Benchmarks

BenchmarkScoreWhat it measuresSourceRetrieved
Intelligence
22.7
Composite score across standardised reasoning, knowledge and problem-solving evaluations, measured independently by Artificial AnalysisSourceSeptember 12, 2026
Coding
49.3
Composite score across standardised coding evaluations, measured independently by Artificial AnalysisSourceSeptember 12, 2026

Gemini 3.5 Flash Lite vs other models

ModelInputOutputContextBest for
$0.39 / 1M tokens
$3.25 / 1M tokens
1.05M tokens
High-volume, cost-efficient agentic tasks
$0.325 / 1M tokens
$1.95 / 1M tokens
1M tokens
Reasoning + agents
$1.95 / 1M tokens
$9.75 / 1M tokens
1.05M tokens
Reasoning + agents
$0.26 / 1M tokens
$1.56 / 1M tokens
1M tokens
Fast, high-volume tasks
$0.975 / 1M tokens
$4.875 / 1M tokens
1.05M tokens
Agentic coding + high-efficiency reasoning

Frequently asked questions

Gemini 3.5 Flash Lite has a 1,048,576 tokens context window and can return up to 65,536 tokens.

Gemini 3.5 Flash Lite takes audio, image, text, video as input and returns audio, text.

Use google/gemini-3.5-flash-lite as the model id. Requests go to https://api.aimlapi.com/v1/chat/completions.

Gemini 3.5 Flash Lite became available on July 21, 2026.

Gemini 3.5 Flash Lite is priced at input $0.39 / 1M tokens, output $3.25 / 1M tokens, cached input $0.039 / 1M tokens.

Yes, Gemini 3.5 Flash Lite can stream responses as they are generated.

Yes, Gemini 3.5 Flash Lite accepts image input alongside text.

Gemini 3.5 Flash Lite was built by Google.

It suits high-volume agentic tasks, translation, and simple data processing where cost efficiency matters.

Not enough information to compare these two models.

Yes, it supports tools, function calling, and parallel tool calls.

Start building with Gemini 3.5 Flash Lite

Get API Key
1000+ models, one API.