Llama 3.1 Nemotron 70B Instruct API

Llama 3.1 Nemotron is an advanced instruction-following language model optimized for high-performance applications.
Output

How to use Llama 3.1 Nemotron 70B Instruct API

Install any OpenAI-compatible SDK, point it at api.aimlapi.com/v1, and set the model to .

OpenAI-compatible — swap the base URL and it works with your existing SDK.

Llama 3.1 Nemotron 70B Instruct API Pricing

TypePrice
Input
Output

Llama 3.1 Nemotron 70B Instruct Benchmarks

BenchmarkScoreWhat it measuresSourceRetrieved
Intelligence
6.9
Composite score across standardised reasoning, knowledge and problem-solving evaluations, measured independently by Artificial AnalysisSourceSeptember 22, 2026
Math
11
Composite score across standardised mathematics evaluations, measured independently by Artificial AnalysisSourceSeptember 22, 2026

Llama 3.1 Nemotron 70B Instruct vs other models

ModelInputOutputContextBest for
$6.5 / 1M tokens
$39 / 1M tokens
1M tokens
Reasoning + agents
$2.6 / 1M tokens
$13 / 1M tokens
1M tokens
Balanced coding + agents
$3.9 / 1M tokens
$19.5 / 1M tokens
1M tokens
Long-context, multimodal & agentic workflows
$1.95 / 1M tokens
$11.7 / 1M tokens
1M tokens
Reasoning + agents

Frequently asked questions

Yes, Llama 3.1 Nemotron 70B Instruct can stream responses as they are generated.

Yes, Llama 3.1 Nemotron 70B Instruct supports both function calling and structured outputs.

Llama 3.1 Nemotron 70B Instruct was built by NVIDIA.

No. Its listed capabilities are function calling, streaming, and structured outputs, with no image input support.

It is accessed via the v1/chat/completions endpoint.

Start building with Llama 3.1 Nemotron 70B Instruct

Get API Key
1000+ models, one API.