import requests r = requests.post( "https://api.aimlapi.com/v1/tts", headers={"Authorization": "Bearer " + AIMLAPI_KEY}, json={ "model": "elevenlabs/eleven_v4", "text": "Hello from AI/ML API", "voice": "Rachel" }, ) print(r.json()["audio"]) # URL to the generated audio
const r = await fetch("https://api.aimlapi.com/v1/tts", { method: "POST", headers: { Authorization: "Bearer " + process.env.AIMLAPI_KEY, "Content-Type": "application/json", }, body: JSON.stringify({ "model": "elevenlabs/eleven_v4", "text": "Hello from AI/ML API", "voice": "Rachel" }), }); console.log((await r.json()).audio);
curl -X POST https://api.aimlapi.com/v1/tts \ -H "Authorization: Bearer $AIMLAPI_KEY" \ -H "Content-Type: application/json" \ -d '{"model":"elevenlabs/eleven_v4","text":"Hello from AI/ML API","voice":"Rachel"}'
OpenAI-compatible — swap the base URL and it works with your existing SDK.
| Type | Price |
|---|---|
| Input | |
| Output | |
Billed per character of input text.
| Benchmark | Score | What it measures | Source | Retrieved |
|---|---|---|---|---|
| TTS Arena | 1321 | Human preference Elo from blind pairwise listening comparisons of speech samples, measured independently by Artificial Analysis | Source | October 6, 2026 |
| Model | Input | Output | Context | Best for |
|---|---|---|---|---|
Eleven v4 This page | Speech synthesis | |||
| Speech synthesis | ||||
| Speech synthesis | ||||
| Speech synthesis |
Eleven v4 is ElevenLabs' text-to-speech model, described by the vendor as its most emotive and highest quality speech synthesis model. It turns text into spoken audio.
$28.6 per 1M characters through AI/ML API. Speech models are billed by characters of input text rather than by tokens.
Text in, audio out. You pass the text and a voice, and the response carries a URL to the generated audio.
Both are ElevenLabs' v4 generation, with the same expressive voices and inline audio tags in 90+ languages. v4 is positioned for highest quality, v4 Turbo for real-time use such as agents and live conversation. Turbo is half the price at $14.3 per 1M characters.
The text-to-speech endpoint, POST /v1/tts, with model set to elevenlabs/eleven_v4.