import requests r = requests.post( "https://api.aimlapi.com/v1/chat/completions", headers={"Authorization": "Bearer " + AIMLAPI_KEY}, json={ "model": "nvidia/llama-3.3-nemotron-super-49b-v1.5", "messages": [ { "role": "user", "content": "Hello!" } ] }, ) print(r.json())
const r = await fetch("https://api.aimlapi.com/v1/chat/completions", { method: "POST", headers: { Authorization: `Bearer ${process.env.AIMLAPI_KEY}`, "Content-Type": "application/json", }, body: JSON.stringify({ "model": "nvidia/llama-3.3-nemotron-super-49b-v1.5", "messages": [ { "role": "user", "content": "Hello!" } ] }), }); console.log(await r.json());
curl -X POST https://api.aimlapi.com/v1/chat/completions \ -H "Authorization: Bearer $AIMLAPI_KEY" \ -H "Content-Type: application/json" \ -d '{"model":"nvidia/llama-3.3-nemotron-super-49b-v1.5","messages":[{"role":"user","content":"Hello!"}]}'
OpenAI-compatible — swap the base URL and it works with your existing SDK.
| Type | Price |
|---|---|
| Input | |
| Output | |
| Benchmark | Score | What it measures | Source | Retrieved |
|---|---|---|---|---|
| MMLU-Pro | 79.5% | Multi-discipline knowledge + reasoning (harder MMLU) | Source | July 8, 2026 |
| LiveCodeBench | 73.6% | Contamination-free competitive programming problems | Source | July 8, 2026 |
| GPQA Diamond | 72% | Google-proof graduate science questions (hardest subset) | Source | July 8, 2026 |
| AIME 2025 | 82.7% | Competition mathematics (AIME), 2025 | Source | July 8, 2026 |
| Model | Input | Output | Context | Best for |
|---|---|---|---|---|
Llama 3.3 Nemotron Super 49B V1.5 This page | Reasoning + agents | |||
| Reasoning + agents | ||||
| Balanced coding + agents | ||||
| Long-context, multimodal & agentic workflows | ||||
| Reasoning + agents |
Llama 3.3 Nemotron Super 49B V1.5 has a 131,072 tokens context window and can return up to 16,384 tokens.
Llama 3.3 Nemotron Super 49B V1.5 takes image, text as input and returns text.
Use nvidia/llama-3.3-nemotron-super-49b-v1.5 as the model id. Requests go to https://api.aimlapi.com/v1/chat/completions.
Llama 3.3 Nemotron Super 49B V1.5 is priced at input $0.52 / 1M, output $0.52 / 1M.
Yes, Llama 3.3 Nemotron Super 49B V1.5 can stream responses as they are generated.
Yes, Llama 3.3 Nemotron Super 49B V1.5 accepts image input alongside text.
Llama 3.3 Nemotron Super 49B V1.5 was built by NVIDIA.
Yes, it is built as a chat model with dedicated reasoning capability.
Yes, it supports tools, parallel tool calls, and function calling.