OpenAI-compatible — swap the base URL and it works with your existing SDK.
| Type | Price |
|---|---|
| Input | |
| Output | |
| Benchmark | Score | What it measures | Source | Retrieved |
|---|---|---|---|---|
| MMMU | 60.3 | College-level multimodal understanding + reasoning | Source | September 22, 2026 |
| ChartQA | 85.5 | Question answering over charts and plots | Source | September 22, 2026 |
| DocVQA | 90.1 | Question answering over document images | Source | September 22, 2026 |
| AI2D | 92.3 | Question answering over science diagrams | Source | September 22, 2026 |
| Model | Input | Output | Context | Best for |
|---|---|---|---|---|
Llama 3.2 90B Vision Instruct Turbo This page | ||||
| Reasoning + agents | ||||
| Balanced coding + agents | ||||
| Long-context, multimodal & agentic workflows | ||||
| Reasoning + agents |
Yes, Llama 3.2 90B Vision Instruct Turbo can stream responses as they are generated.
Yes, Llama 3.2 90B Vision Instruct Turbo supports both function calling and structured outputs.
Llama 3.2 90B Vision Instruct Turbo was built by Meta.
It is accessed via the v1/chat/completions endpoint.