Llama 3.2 90B Vision Instruct Turbo API

Powerful multimodal AI model for advanced visual and language processing tasks.
Output

How to use Llama 3.2 90B Vision Instruct Turbo API

Install any OpenAI-compatible SDK, point it at api.aimlapi.com/v1, and set the model to .

OpenAI-compatible — swap the base URL and it works with your existing SDK.

Llama 3.2 90B Vision Instruct Turbo API Pricing

TypePrice
Input
Output

Llama 3.2 90B Vision Instruct Turbo Benchmarks

BenchmarkScoreWhat it measuresSourceRetrieved
MMMU
60.3
College-level multimodal understanding + reasoningSourceSeptember 22, 2026
ChartQA
85.5
Question answering over charts and plotsSourceSeptember 22, 2026
DocVQA
90.1
Question answering over document imagesSourceSeptember 22, 2026
AI2D
92.3
Question answering over science diagramsSourceSeptember 22, 2026

Llama 3.2 90B Vision Instruct Turbo vs other models

ModelInputOutputContextBest for
$6.5 / 1M tokens
$39 / 1M tokens
1M tokens
Reasoning + agents
$2.6 / 1M tokens
$13 / 1M tokens
1M tokens
Balanced coding + agents
$3.9 / 1M tokens
$19.5 / 1M tokens
1M tokens
Long-context, multimodal & agentic workflows
$1.95 / 1M tokens
$11.7 / 1M tokens
1M tokens
Reasoning + agents

Frequently asked questions

Yes, Llama 3.2 90B Vision Instruct Turbo can stream responses as they are generated.

Yes, Llama 3.2 90B Vision Instruct Turbo supports both function calling and structured outputs.

Llama 3.2 90B Vision Instruct Turbo was built by Meta.

It is accessed via the v1/chat/completions endpoint.

Start building with Llama 3.2 90B Vision Instruct Turbo

Get API Key
1000+ models, one API.