OpenAI-compatible — swap the base URL and it works with your existing SDK.
| Type | Price |
|---|---|
| Input | |
| Output | |
| Benchmark | Score | What it measures | Source | Retrieved |
|---|---|---|---|---|
| Intelligence | 6.9 | Composite score across standardised reasoning, knowledge and problem-solving evaluations, measured independently by Artificial Analysis | Source | September 22, 2026 |
| Math | 11 | Composite score across standardised mathematics evaluations, measured independently by Artificial Analysis | Source | September 22, 2026 |
| Model | Input | Output | Context | Best for |
|---|---|---|---|---|
Llama 3.1 Nemotron 70B Instruct This page | ||||
| Reasoning + agents | ||||
| Balanced coding + agents | ||||
| Long-context, multimodal & agentic workflows | ||||
| Reasoning + agents |
Yes, Llama 3.1 Nemotron 70B Instruct can stream responses as they are generated.
Yes, Llama 3.1 Nemotron 70B Instruct supports both function calling and structured outputs.
Llama 3.1 Nemotron 70B Instruct was built by NVIDIA.
No. Its listed capabilities are function calling, streaming, and structured outputs, with no image input support.
It is accessed via the v1/chat/completions endpoint.