import requests r = requests.post( "https://api.aimlapi.com/v1/chat/completions", headers={"Authorization": "Bearer " + AIMLAPI_KEY}, json={ "model": "nvidia/nemotron-3.5-lightning", "messages": [ { "role": "user", "content": "Hello!" } ] }, ) print(r.json())
const r = await fetch("https://api.aimlapi.com/v1/chat/completions", { method: "POST", headers: { Authorization: `Bearer ${process.env.AIMLAPI_KEY}`, "Content-Type": "application/json", }, body: JSON.stringify({ "model": "nvidia/nemotron-3.5-lightning", "messages": [ { "role": "user", "content": "Hello!" } ] }), }); console.log(await r.json());
curl -X POST https://api.aimlapi.com/v1/chat/completions \ -H "Authorization: Bearer $AIMLAPI_KEY" \ -H "Content-Type: application/json" \ -d '{"model":"nvidia/nemotron-3.5-lightning","messages":[{"role":"user","content":"Hello!"}]}'
OpenAI-compatible — swap the base URL and it works with your existing SDK.
| Type | Price |
|---|---|
| Input | |
| Output | |
This model is currently free to use.
| Benchmark | Score | What it measures | Source | Retrieved |
|---|---|---|---|---|
| Intelligence | 13.6 | Composite score across standardised reasoning, knowledge and problem-solving evaluations, measured independently by Artificial Analysis | Source | September 12, 2026 |
| Coding | 26.8 | Composite score across standardised coding evaluations, measured independently by Artificial Analysis | Source | September 12, 2026 |
| SWE-bench Verified | 51.56 | Resolving verified real GitHub issues | Source | September 21, 2026 |
| Terminal-Bench | 24.58 | Autonomous shell/terminal task completion | Source | September 21, 2026 |
| MMLU-Pro | 81.94 | Multi-discipline knowledge + reasoning (harder MMLU) | Source | September 21, 2026 |
| GPQA Diamond | 75.44 | Google-proof graduate science questions (hardest subset) | Source | September 21, 2026 |
| Humanity's Last Exam | 11.72 | Expert-level questions across many domains | Source | September 21, 2026 |
| Model | Input | Output | Context | Best for |
|---|---|---|---|---|
Nemotron 3.5 Lightning This page | High-throughput agentic workloads | |||
| Long-context, multimodal & agentic workflows | ||||
| Reasoning + agents | ||||
| Complex reasoning, coding and agentic workflows | ||||
| Fast, low-cost reasoning and tool use at scale |
Nemotron 3.5 Lightning has a 1,000,000 tokens context window and can return up to 65,536 tokens.
Nemotron 3.5 Lightning takes text as input and returns text.
Use nvidia/nemotron-3.5-lightning as the model id. Requests go to https://api.aimlapi.com/v1/chat/completions.
Nemotron 3.5 Lightning became available on August 11, 2026.
Nemotron 3.5 Lightning is priced at input $0 / 1M, output $0 / 1M.
Yes, Nemotron 3.5 Lightning can stream responses as they are generated.
Nemotron 3.5 Lightning was built by NVIDIA.
It is a mixture-of-experts model with 30B total parameters, using 3B active parameters per task.