import requests r = requests.post( "https://api.aimlapi.com/v1/chat/completions", headers={"Authorization": "Bearer " + AIMLAPI_KEY}, json={ "model": "google/gemini-4-argon", "messages": [ { "role": "user", "content": "Hello!" } ] }, ) print(r.json())
const r = await fetch("https://api.aimlapi.com/v1/chat/completions", { method: "POST", headers: { Authorization: `Bearer ${process.env.AIMLAPI_KEY}`, "Content-Type": "application/json", }, body: JSON.stringify({ "model": "google/gemini-4-argon", "messages": [ { "role": "user", "content": "Hello!" } ] }), }); console.log(await r.json());
curl -X POST https://api.aimlapi.com/v1/chat/completions \ -H "Authorization: Bearer $AIMLAPI_KEY" \ -H "Content-Type: application/json" \ -d '{"model":"google/gemini-4-argon","messages":[{"role":"user","content":"Hello!"}]}'
OpenAI-compatible — swap the base URL and it works with your existing SDK.
| Type | Price |
|---|---|
| Input | |
| Output | |
| Cached input |
Coming soon — not yet available on AI/ML API. Prices are listed ahead of launch and are not billable until the model goes live.
| Benchmark | Score | What it measures | Source | Retrieved |
|---|---|---|---|---|
| Intelligence | 52.6 | Composite score across standardised reasoning, knowledge and problem-solving evaluations, measured independently by Artificial Analysis | Source | September 30, 2026 |
Figures reported by Google; not independently reproduced.
Not yet. Gemini 4 Argon is listed ahead of launch and cannot be called on AI/ML API today. This page is updated as soon as access opens, and the prices shown here are not billable until then.
Up to 1,000,000 tokens. That is the headline change in this release: the previous ceiling was 64,000 tokens, so a generation that used to be split across many calls can now come back in a single response.
On AI/ML API, Gemini 4 Argon is priced at $2.60 per 1M input tokens and $13.00 per 1M output tokens, with cached input at $0.13 per 1M tokens. Billing starts only once the model is live.
Google has not published an input context window for Argon yet. The 1,000,000-token figure announced with the model is its output limit, not its input window, so this page leaves the context window blank rather than carrying over a number from an earlier Gemini model. It will be filled in once Google states it.
Google positions Argon for agentic software engineering and long-horizon work: it cites agents analysing fleet-wide profiling telemetry and migrating C/C++ codebases. The raised output ceiling also suits long-form generation, where a whole document, report or code change is returned in one response.
At announcement Argon is in limited release. Google is rolling it out first to a set of trusted cyber defenders through its Fairwind Program, and says broader access will follow for paid API customers and Google AI Ultra subscribers. On AI/ML API it becomes callable with a single model ID once that access opens.
No. Google reports Argon ahead of competing models from OpenAI and Anthropic on several benchmarks, including long-video understanding. Those results are vendor-reported and have not been independently reproduced.