Nemotron 3.5 Lightning API

nvidia/nemotron-3.5-lightning
NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model with 3B active parameters out of 30B total, built for high-throughput agentic workloads and specialized task execution with up to 1M context length.
Context
1M tokens
Input
$0 / 1M
Output
$0 / 1M
Released
Aug 11, 2026

How to use Nemotron 3.5 Lightning API

Install any OpenAI-compatible SDK, point it at api.aimlapi.com/v1, and set the model to nvidia/nemotron-3.5-lightning.
import requests

r = requests.post(
    "https://api.aimlapi.com/v1/chat/completions",
    headers={"Authorization": "Bearer " + AIMLAPI_KEY},
    json={
      "model": "nvidia/nemotron-3.5-lightning",
      "messages": [
        {
          "role": "user",
          "content": "Hello!"
        }
      ]
    },
)
print(r.json())
const r = await fetch("https://api.aimlapi.com/v1/chat/completions", {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.AIMLAPI_KEY}`,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
    "model": "nvidia/nemotron-3.5-lightning",
    "messages": [
      {
        "role": "user",
        "content": "Hello!"
      }
    ]
  }),
});
console.log(await r.json());
curl -X POST https://api.aimlapi.com/v1/chat/completions \
  -H "Authorization: Bearer $AIMLAPI_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"nvidia/nemotron-3.5-lightning","messages":[{"role":"user","content":"Hello!"}]}'

OpenAI-compatible — swap the base URL and it works with your existing SDK.

Nemotron 3.5 Lightning API Pricing

TypePrice
Input
$0 / 1M
Output
$0 / 1M

This model is currently free to use.

Nemotron 3.5 Lightning Benchmarks

BenchmarkScoreWhat it measuresSourceRetrieved
Intelligence
13.6
Composite score across standardised reasoning, knowledge and problem-solving evaluations, measured independently by Artificial AnalysisSourceSeptember 12, 2026
Coding
26.8
Composite score across standardised coding evaluations, measured independently by Artificial AnalysisSourceSeptember 12, 2026
SWE-bench Verified
51.56
Resolving verified real GitHub issuesSourceSeptember 21, 2026
Terminal-Bench
24.58
Autonomous shell/terminal task completionSourceSeptember 21, 2026
MMLU-Pro
81.94
Multi-discipline knowledge + reasoning (harder MMLU)SourceSeptember 21, 2026
GPQA Diamond
75.44
Google-proof graduate science questions (hardest subset)SourceSeptember 21, 2026
Humanity's Last Exam
11.72
Expert-level questions across many domainsSourceSeptember 21, 2026

Nemotron 3.5 Lightning vs other models

ModelInputOutputContextBest for
$0 / 1M
$0 / 1M
1M tokens
High-throughput agentic workloads
$3.9 / 1M tokens
$19.5 / 1M tokens
1M tokens
Long-context, multimodal & agentic workflows
$2.0631 / 1M tokens
$12.3786 / 1M tokens
400K tokens
Reasoning + agents
$2.6 / 1M tokens
$7.8 / 1M tokens
1M tokens
Complex reasoning, coding and agentic workflows
$0.103155 / 1M tokens
$0.302588 / 1M tokens
256K tokens
Fast, low-cost reasoning and tool use at scale

Frequently asked questions

Nemotron 3.5 Lightning has a 1,000,000 tokens context window and can return up to 65,536 tokens.

Nemotron 3.5 Lightning takes text as input and returns text.

Use nvidia/nemotron-3.5-lightning as the model id. Requests go to https://api.aimlapi.com/v1/chat/completions.

Nemotron 3.5 Lightning became available on August 11, 2026.

Nemotron 3.5 Lightning is priced at input $0 / 1M, output $0 / 1M.

Yes, Nemotron 3.5 Lightning can stream responses as they are generated.

Nemotron 3.5 Lightning was built by NVIDIA.

It is a mixture-of-experts model with 30B total parameters, using 3B active parameters per task.

Start building with Nemotron 3.5 Lightning

Get API Key
1000+ models, one API.