Drafted with AI tools and edited by a human. The figures and conclusions were checked by a GPUniq editor.
TL;DR
Eleven v4 is the latest flagship text-to-speech model from ElevenLabs, available via Runway Dev. It supports 90+ languages, inline delivery tags, and natural language prompts. The low-latency Eleven v4 Turbo variant reaches ~100 ms median inference latency. API pricing runs at a promotional rate of 2.2 credits for the first 1,000 characters until October 12, 2026 PT, after which standard pricing is 5 credits per 1,000 characters.
What Features Power Eleven v4?
Eleven v4 upgrades base voice cloning and context handling. The architecture natively processes over 90 languages. Professional Voice Cloning (PVC) requires verified consent before training, and clones created prior to v4 must be retrained to work effectively with v4.
The model accepts direct natural language direction prompts alongside input text. It also reads inline audio delivery tags to alter emotional delivery during generation.
[whispers] Keep your voice down. [pause] [laughs] They might hear us!
According to the ElevenLabs Blog, Eleven v4 achieved a ~75% listener preference score in blind head-to-head tests against competing models like Cartesia Sonic 3.6 and Google Gemini 3.8 Flash-Lite. To maintain speech continuity across calls, pass preceding and following text strings into the context parameters.
How Much Does Eleven v4 Cost?
API generation costs use a character-based credit system. Promotional pricing applies across all tiers, including Free accounts.
| Generation Model | Promo Rate (Until Oct 12, 2026 PT) | Standard Rate (Post Oct 12, 2026 PT) |
|---|---|---|
| Eleven v4 | 2.2 credits (first 1,000 characters) | 5 credits per 1,000 characters |
Standard usage costs 5 credits per 1,000 characters after the promotional period ends.
How Fast Is Eleven v4 Turbo?
According to the ElevenLabs Blog, Eleven v4 Turbo hits a median inference latency of ~100 ms.
| Metric | Target |
|---|---|
| Eleven v4 Turbo Latency | ~100 ms |
Output quality is tuned heavily for expressive content. Because of this tuning, the low-latency Turbo variant favors pure speed over absolute audio fidelity.
How Does Eleven v4 Compare?
Choosing between text-to-speech models comes down to your latency and fidelity needs.
| Metric | Eleven v4 | Eleven v4 Turbo | Cartesia Sonic 3.6 / Gemini 3.8 Flash-Lite |
|---|---|---|---|
| Primary Focus | Expressive speech & fidelity | Real-time response | Competitor baseline |
| Median Latency | - | ~100 ms | - |
| Language Count | 90+ native languages | 90+ native languages | - |
Eleven v4 is an upgrade over its predecessor, Eleven v3. Against Cartesia Sonic 3.6 and Google Gemini 3.8 Flash-Lite, Eleven v4 won ~75% of blind listener evaluation tests. Use Turbo for live agents where 100 ms latency matters. Use standard v4 for expressive audio output.
What Are the API Constraints?
Eleven v4 sets strict boundaries on generation payloads:
- Hard character limit: 10,000 characters per single payload (roughly 10 minutes of output audio).
- Pre-v4 voice clones: Retrain legacy Professional Voice Clones to run them on v4 endpoints.
- Model trade-offs: Turbo prioritizes latency over absolute audio fidelity.
Chunk long inputs into separate requests to stay under the limit. Use preceding and following context parameters across consecutive chunks to maintain speech continuity.
How Do You Deploy v4 via Runway Dev?
Run Eleven v4 inside Runway Dev by calling the /v1/text_to_speech endpoint with model_id set to eleven_v4.
import requests
url = "https://api.dev.runwayml.com/v1/text_to_speech"
headers = {
"Authorization": "Bearer YOUR_RUNWAY_API_KEY",
"Content-Type": "application/json"
}
payload = {
"model_id": "eleven_v4",
"text": "[whispers] Initializing system sequence. [pause] All systems nominal.",
"voice_settings": {
"stability": 0.5,
"similarity_boost": 0.75
}
}
response = requests.post(url, json=payload, headers=headers)
if response.status_code == 402:
raise Exception("Credit quota exhausted. Check your balance.")
elif response.status_code != 200:
raise Exception(f"API Error {response.status_code}: {response.text}")
with open("output.mp3", "wb") as f:
f.write(response.content)
Handle API errors cleanly in production code. Catch API status codes to intercept exhausted credit balances before retrying failed requests. Switch your payload target to eleven_v4_turbo whenever you build live conversational voice bots.
Related on the GPUniq blog
Methodology
Facts are compiled from the vendor's own announcement and documentation, linked below, at the time of writing; where the vendor hasn't published a number it is marked as such. GPUniq prices, when the model is already available here, come from our live catalog.
Sources
- 1.API Changelog & Updates | Runway Dev — Eleven v4 on Runway Dev — Runway Dev (accessed )
- 2.One API for AI Video, Image and Audio | Runway Dev — text_to_speech endpoint for eleven_v4 — Runway Dev (accessed )
- 3.Eleven v4: Our most expressive text-to-speech model yet — ElevenLabs (accessed )
- 4.Models | ElevenLabs Documentation — v4 overview — ElevenLabs (accessed )
Want cheap GPUs for your next project?
Browse live GPU prices and rent the right card in seconds — H100, A100, RTX 4090, and 50+ more models.
