Nemotron 3 Nano 30B A3B API

nvidia/nemotron-3-nano-30b-a3b
A sparse mixture-of-experts language model that activates just 3B parameters per token while drawing on 30B of learned knowledge, built from the ground up for agentic AI systems, production RAG pipelines, and long-context reasoning at scale.
Context
256K tokens
Input
$0.06877 / 1M tokens
Output
$0.27508 / 1M tokens
Released
Apr 30, 2026

How to use Nemotron 3 Nano 30B A3B API

Install any OpenAI-compatible SDK, point it at api.aimlapi.com/v1, and set the model to nvidia/nemotron-3-nano-30b-a3b.
import requests

r = requests.post(
    "https://api.aimlapi.com/v1/chat/completions",
    headers={"Authorization": "Bearer " + AIMLAPI_KEY},
    json={
      "model": "nvidia/nemotron-3-nano-30b-a3b",
      "messages": [
        {
          "role": "user",
          "content": "Hello!"
        }
      ]
    },
)
print(r.json())
const r = await fetch("https://api.aimlapi.com/v1/chat/completions", {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.AIMLAPI_KEY}`,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
    "model": "nvidia/nemotron-3-nano-30b-a3b",
    "messages": [
      {
        "role": "user",
        "content": "Hello!"
      }
    ]
  }),
});
console.log(await r.json());
curl -X POST https://api.aimlapi.com/v1/chat/completions \
  -H "Authorization: Bearer $AIMLAPI_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"nvidia/nemotron-3-nano-30b-a3b","messages":[{"role":"user","content":"Hello!"}]}'

OpenAI-compatible — swap the base URL and it works with your existing SDK.

Nemotron 3 Nano 30B A3B API Pricing

TypePrice
Input
$0.06877 / 1M tokens
Output
$0.27508 / 1M tokens

Nemotron 3 Nano 30B A3B Benchmarks

BenchmarkScoreWhat it measuresSourceRetrieved
Humanity's Last Exam
10.6%
Expert-level questions across many domainsSourceJuly 13, 2026
MMLU-Pro
78.3%
Multi-discipline knowledge + reasoning (harder MMLU)SourceJuly 13, 2026
LiveCodeBench
68.3%
Contamination-free competitive programming problemsSourceJuly 13, 2026
GPQA Diamond
73.0%
Google-proof graduate science questions (hardest subset)SourceJuly 13, 2026
AIME 2025
89.1%
Competition mathematics (AIME), 2025SourceJuly 13, 2026
Intelligence
8.9
Composite score across standardised reasoning, knowledge and problem-solving evaluations, measured independently by Artificial AnalysisSourceSeptember 22, 2026
Coding
14.4
Composite score across standardised coding evaluations, measured independently by Artificial AnalysisSourceSeptember 22, 2026
Math
91
Composite score across standardised mathematics evaluations, measured independently by Artificial AnalysisSourceSeptember 22, 2026

Nemotron 3 Nano 30B A3B vs other models

ModelInputOutputContextBest for
$0.06877 / 1M tokens
$0.27508 / 1M tokens
256K tokens
Coding + agents
$6.5 / 1M tokens
$39 / 1M tokens
1M tokens
Reasoning + agents
$2.6 / 1M tokens
$13 / 1M tokens
1M tokens
Balanced coding + agents
$3.9 / 1M tokens
$19.5 / 1M tokens
1M tokens
Long-context, multimodal & agentic workflows
$0.65 / 1M tokens
$3.9 / 1M tokens
1M tokens
Reasoning + agents

Frequently asked questions

Nemotron 3 Nano 30B A3B has a 262,144 tokens context window and can return up to 228,000 tokens.

Nemotron 3 Nano 30B A3B takes text as input and returns text.

Use nvidia/nemotron-3-nano-30b-a3b as the model id. Requests go to https://api.aimlapi.com/v1/chat/completions.

Nemotron 3 Nano 30B A3B became available on April 30, 2026.

Nemotron 3 Nano 30B A3B is priced at input $0.06877 / 1M tokens, output $0.27508 / 1M tokens.

Yes, Nemotron 3 Nano 30B A3B can stream responses as they are generated.

Nemotron 3 Nano 30B A3B was built by NVIDIA.

Start building with Nemotron 3 Nano 30B A3B

Get API Key
1000+ models, one API.