Llama 3.3 Nemotron Super 49B V1.5 API

nvidia/llama-3.3-nemotron-super-49b-v1.5
Llama 3.3 Nemotron Super 49B V1.5 on AIMLAPI.
Context
131K tokens
Input
$0.52 / 1M
Output
$0.52 / 1M

How to use Llama 3.3 Nemotron Super 49B V1.5 API

Install any OpenAI-compatible SDK, point it at api.aimlapi.com/v1, and set the model to nvidia/llama-3.3-nemotron-super-49b-v1.5.
import requests

r = requests.post(
    "https://api.aimlapi.com/v1/chat/completions",
    headers={"Authorization": "Bearer " + AIMLAPI_KEY},
    json={
      "model": "nvidia/llama-3.3-nemotron-super-49b-v1.5",
      "messages": [
        {
          "role": "user",
          "content": "Hello!"
        }
      ]
    },
)
print(r.json())
const r = await fetch("https://api.aimlapi.com/v1/chat/completions", {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.AIMLAPI_KEY}`,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
    "model": "nvidia/llama-3.3-nemotron-super-49b-v1.5",
    "messages": [
      {
        "role": "user",
        "content": "Hello!"
      }
    ]
  }),
});
console.log(await r.json());
curl -X POST https://api.aimlapi.com/v1/chat/completions \
  -H "Authorization: Bearer $AIMLAPI_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"nvidia/llama-3.3-nemotron-super-49b-v1.5","messages":[{"role":"user","content":"Hello!"}]}'

OpenAI-compatible — swap the base URL and it works with your existing SDK.

Llama 3.3 Nemotron Super 49B V1.5 API Pricing

TypePrice
Input
$0.52 / 1M
Output
$0.52 / 1M

Llama 3.3 Nemotron Super 49B V1.5 Benchmarks

BenchmarkScoreWhat it measuresSourceRetrieved
MMLU-Pro
79.5%
Multi-discipline knowledge + reasoning (harder MMLU)SourceJuly 8, 2026
LiveCodeBench
73.6%
Contamination-free competitive programming problemsSourceJuly 8, 2026
GPQA Diamond
72%
Google-proof graduate science questions (hardest subset)SourceJuly 8, 2026
AIME 2025
82.7%
Competition mathematics (AIME), 2025SourceJuly 8, 2026

Llama 3.3 Nemotron Super 49B V1.5 vs other models

ModelInputOutputContextBest for
$0.52 / 1M
$0.52 / 1M
131K tokens
Reasoning + agents
$6.5 / 1M tokens
$39 / 1M tokens
1.05M tokens
Reasoning + agents
$2.6 / 1M tokens
$13 / 1M tokens
1M tokens
Balanced coding + agents
$3.9 / 1M tokens
$19.5 / 1M tokens
1M tokens
Long-context, multimodal & agentic workflows
$0.65 / 1M tokens
$3.9 / 1M tokens
1.05M tokens
Reasoning + agents

Frequently asked questions

Llama 3.3 Nemotron Super 49B V1.5 has a 131,072 tokens context window and can return up to 16,384 tokens.

Llama 3.3 Nemotron Super 49B V1.5 takes image, text as input and returns text.

Use nvidia/llama-3.3-nemotron-super-49b-v1.5 as the model id. Requests go to https://api.aimlapi.com/v1/chat/completions.

Llama 3.3 Nemotron Super 49B V1.5 is priced at input $0.52 / 1M, output $0.52 / 1M.

Yes, Llama 3.3 Nemotron Super 49B V1.5 can stream responses as they are generated.

Yes, Llama 3.3 Nemotron Super 49B V1.5 accepts image input alongside text.

Llama 3.3 Nemotron Super 49B V1.5 was built by NVIDIA.

Yes, it is built as a chat model with dedicated reasoning capability.

Yes, it supports tools, parallel tool calls, and function calling.

Start building with Llama 3.3 Nemotron Super 49B V1.5

Get API Key
1000+ models, one API.