Gemini 3.8 Flash API

google/gemini-3.8-flash
Gemini 3.8 Flash is Google's most intelligent Flash model, built for long-horizon software engineering, autonomous agents, and complex enterprise workflows.
Context
1.05M tokens
Input
$0.975 / 1M tokens
Output
$4.875 / 1M tokens
Released
Sep 2, 2026

How to use Gemini 3.8 Flash API

Install any OpenAI-compatible SDK, point it at api.aimlapi.com/v1, and set the model to google/gemini-3.8-flash.
import requests

r = requests.post(
    "https://api.aimlapi.com/v1/chat/completions",
    headers={"Authorization": "Bearer " + AIMLAPI_KEY},
    json={
      "model": "google/gemini-3.8-flash",
      "messages": [
        {
          "role": "user",
          "content": "Hello!"
        }
      ]
    },
)
print(r.json())
const r = await fetch("https://api.aimlapi.com/v1/chat/completions", {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.AIMLAPI_KEY}`,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
    "model": "google/gemini-3.8-flash",
    "messages": [
      {
        "role": "user",
        "content": "Hello!"
      }
    ]
  }),
});
console.log(await r.json());
curl -X POST https://api.aimlapi.com/v1/chat/completions \
  -H "Authorization: Bearer $AIMLAPI_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"google/gemini-3.8-flash","messages":[{"role":"user","content":"Hello!"}]}'

OpenAI-compatible — swap the base URL and it works with your existing SDK.

Gemini 3.8 Flash API Pricing

TypePrice
Input
$0.975 / 1M tokens
Output
$4.875 / 1M tokens
Cached input
$0.0975 / 1M tokens

Google Search grounding billed separately at $0.0455 per call. Cached input priced at $0.0975/1M.

Gemini 3.8 Flash Benchmarks

BenchmarkScoreWhat it measuresSourceRetrieved
Intelligence
41.2
Composite score across standardised reasoning, knowledge and problem-solving evaluations, measured independently by Artificial AnalysisSourceSeptember 12, 2026
Coding
76.3
Composite score across standardised coding evaluations, measured independently by Artificial AnalysisSourceSeptember 12, 2026

Gemini 3.8 Flash vs other models

ModelInputOutputContextBest for
$0.975 / 1M tokens
$4.875 / 1M tokens
1.05M tokens
Agentic coding + high-efficiency reasoning
$0.975 / 1M tokens
$4.875 / 1M tokens
1.05M tokens
Agentic coding + high-efficiency reasoning
$1.95 / 1M tokens
$9.75 / 1M tokens
1.05M tokens
Reasoning + agents
$2.6 / 1M tokens
$7.8 / 1M tokens
500K tokens
Reasoning + agents with long context
$5.2 / 1M tokens
$26 / 1M tokens
1M tokens
Frontier reasoning + agents
$0.39 / 1M tokens
$3.25 / 1M tokens
1.05M tokens
High-volume, cost-efficient agentic tasks

Frequently asked questions

Gemini 3.8 Flash has a 1,048,576 tokens context window and can return up to 65,536 tokens.

Gemini 3.8 Flash takes image, text as input and returns text.

Use google/gemini-3.8-flash as the model id. Requests go to https://api.aimlapi.com/v1/chat/completions.

Gemini 3.8 Flash became available on September 2, 2026.

Gemini 3.8 Flash is priced at input $0.975 / 1M tokens, output $4.875 / 1M tokens, cached input $0.0975 / 1M tokens.

Yes, Gemini 3.8 Flash can stream responses as they are generated.

Yes, Gemini 3.8 Flash accepts image input alongside text.

Gemini 3.8 Flash was built by Google.

Start building with Gemini 3.8 Flash

Get API Key
1000+ models, one API.