ElevenLabs Eleven v4: Voice generation levels up with Turbo and real-time latency
🔎 Two models, two fronts: quality and the millisecond
On September 28, 2026, ElevenLabs launched Eleven v4 and Eleven v4 Turbo, formalizing a two-tier strategy: v4 for studio quality, Turbo for real-time latency in conversational agents. The launch paper is co-signed by both founders, Mati Staniszewski and Piotr Dabkowski — a strong signal about the ambition behind this update.
Why now? Because the TTS throne has never been so contested. Cartesia Sonic 3.6 is attacking the agent market head-on, Google is slashing prices with Gemini 3.8 Flash TTS, and open source — with VoiceStudio at the forefront, boasting 646 languages and 24,500 stars on GitHub — offers an escape route for teams allergic to proprietary APIs.
My analysis, with the numbers under the microscope: v4 consolidates the quality crown, but Turbo is the real strategic breakthrough. 150 ms before the first syllable is below the threshold of a human pause. The friction point remains pricing — $80 per million characters, the most expensive of the top trio. Breakdown.
The essentials
- Quality: Eleven v4 is #1 on Artificial Analysis's Provider Voice TTS Arena with an Elo of 1,319 (1,674 appearances, September 2026), ahead of Cartesia Sonic 3.6 (1,276) and Gemini 3.8 Flash TTS (1,267).
- Real-time: Eleven v4 Turbo shows ~100 ms median inference latency and 150 ms time-to-first-speech in independent WebSocket benchmarks, versus 262 ms for Sonic 3.6 and 814 ms for GPT-4o mini TTS.
- Expressiveness: inline tags for emotions, laughter, sighs, accents and SFX, multi-speaker dialogues, ~75% listener preference in blind evaluations (vendor figure).
- Languages: 90+ languages versus 70+ for Eleven v3, voice identity transfer across languages, instant cloning from 10 seconds of audio.
- Pricing: $80/1M characters for v4 (September 2026) — 63% more expensive than Sonic 3.6 ($49) and nearly 5x Gemini 3.8 Flash TTS ($16.49).
- Integration: v4 ↔ Turbo switch via a single
model_idparameter (eleven_v4/eleven_v4_turbo), two weeks of free generations and 2x credits for Creator+ plans.
Recommended Tools
| Tool | Main use | Price (September 2026) | Ideal for |
|---|---|---|---|
| ElevenLabs Eleven v4 | Studio TTS, narration, high-fidelity voice cloning | $80/1M characters (check on elevenlabs.io) | Audiobooks, brand voices, dubbing |
| Eleven v4 Turbo | Real-time TTS for agents | Pricing not detailed at launch (check on elevenlabs.io) | Voice agents, support, telephony |
| Cartesia Sonic 3.6 | Real-time TTS | $49/1M characters (check on cartesia.ai) | Budget-conscious agents |
| Gemini 3.8 Flash TTS | Low-cost TTS | $16.49/1M characters (check on ai.google.dev) | High volumes, prototyping |
| VoiceStudio (open source) | Self-hosted TTS | Free (GPU costs at your expense) | Data sovereignty, very high volumes |
Pricing shifts every quarter in this market: always check official prices before budgeting.
So what exactly is Eleven v4?
Eleven v4 is the most expressive and emotional speech synthesis model ever released by ElevenLabs, live since September 28, 2026 in the creator app, the agents platform, and the developer API. Where v3 read a text, v4 performs a script.
Performance direction is driven with inline tags that control emotion, pacing, reactions, sound effects, and style. The model can laugh, sigh, and switch accents — including playing an angry French speaker, a detail that matters for localization. Multi-speaker dialogues are handled natively, which simplifies audio fiction production.
On the raw performance side, the progress is measurable: 73.4 characters/s generated versus 42.5 for v3 (+73%), and 90+ languages versus 70+. Pronunciation robustness reaches 91.7%, the best score ever measured by Artificial Analysis, versus 85.6% for v3. For French, a language of liaisons and exceptions, that's a concrete argument.
In blind evaluations, v4 earns roughly 75% listener preference according to the launch paper. That's a vendor-published figure, to be taken with the usual caution — but independent leaderboards confirm the ranking, as we'll see below.
Voice cloning shifts into a higher gear
Instant Voice Clone produces a high-fidelity clone from just 10 seconds of audio. The barrier to entry for cloning is no longer technical or financial: it is now ethical and legal, a point I come back to in the common mistakes section.
Speaker consistency over long generations — a long-standing weak point of voice models — is explicitly highlighted by ElevenLabs on both v4 and Turbo.
Another underrated capability: cross-language transfer. A voice recorded in one language can speak another while retaining the speaker's identity and adopting a native accent. For a brand rolling out the same voice ambassador across ten markets, the production gain is substantial.
Turbo: 100 ms, the latency that truly unlocks voice agents
Eleven v4 Turbo is the real-time variant of v4, co-optimized with ElevenAgents, with a median inference latency of around 100 ms — faster than the natural pause between two humans responding to each other. This is the number that changes the game for agents.
The real architectural innovation is bidirectional streaming over WebSockets.
Bidirectional streaming, the silent revolution
Concretely: the client sends text as the LLM generates it and receives audio before the sentence is even finished. Synthesis starts while the language model is still thinking, and perceived latency approaches that of TTS alone.
This design eliminates the classic pattern of "the LLM finishes its sentence, then the TTS starts," which stacked two waits on top of each other. For developers, this means a simpler agent architecture and a conversation that no longer "clunks" between turns of speech.
On independent WebSocket benchmarks, Turbo reaches a median time-to-first-speech of 150 ms. The hierarchy is unequivocal:
| Model | Median time-to-first-speech (WebSocket, September 2026) |
|---|---|
| Eleven v4 Turbo | 150 ms |
| Cartesia Sonic 3.6 | 262 ms |
| OpenAI GPT-4o mini TTS | 814 ms |
814 ms is unusable in conversation: the other person perceives a gap, hesitates, cuts themselves off. At 150 ms, we're back below the human pause — turn-taking research (Stivers et al., PNAS, 2009) places the average gap between two speakers at around 200 ms. OpenAI, outpaced with its classic TTS, is betting on a dedicated real-time line: we covered GPT-Realtime-2 and its three voice models in a separate article.
Turbo also retains v4's expressive controls and supports Professional Voice Clone throughout a call, as confirmed by TestingCatalog. Your agent keeps your brand's voice from the first word to the last, script-directed laughs and pauses included.
Facing the Competition: What the Numbers Really Say
On quality, Eleven v4 dominates the independent rankings — but not all of them, and that's where the analysis gets interesting. The consolidated data from BenchLM and the voice leaderboard from Artificial Analysis converge on three tables:
| Ranking (Artificial Analysis, September 2026) | Eleven v4 Position | Detail |
|---|---|---|
| Provider Voice TTS Arena | #1 | Elo 1,319 over 1,674 appearances, ahead of Sonic 3.6 (1,276) and Gemini 3.8 Flash TTS (1,267) |
| Controlled Voice | #2 | Elo 1,157, behind Alibaba Qwen-Audio-3.1-TTS-Plus (1,178), well ahead of Eleven v3 (1,073) |
| Pronunciation Robustness | #1 (record) | 91.7%, ahead of Gemini 3.8 Flash TTS (89.5%) and Gemini 3.1 Flash TTS (88.2%) |
v4 is also #1 in all four measured categories of the arena: Customer Service, Assistants, Knowledge Sharing, and Entertainment. The Elo gap with Cartesia (43 points) remains moderate, however: the hierarchy is clear, but the dominance isn't overwhelming.
The most instructive point is Controlled Voice: when performance is fully directed, Alibaba Qwen-Audio-3.1-TTS-Plus takes first place. ElevenLabs dominates the open arena, not fine-grained control — a nuance that matters for dubbing studios, far less for conversational agents.
What About Open Source? The VoiceStudio Case
VoiceStudio, an open source project with 24,500 stars on GitHub, covers 646 languages — far more than ElevenLabs' 90+. On paper, the gap is humiliating. In practice, this figure should be read for what it is: broad language coverage says nothing about per-language quality or pronunciation robustness, measured at 91.7% on the v4 side.
My position is unambiguous: VoiceStudio is the rational option if data sovereignty is non-negotiable — healthcare, finance, public sector — or if your volume makes the API prohibitive. You then trade simplicity for infrastructure: GPUs to size, latency to tune yourself, no SLA, no turnkey agent integration.
To evaluate the open source option without committing a heavy budget, a GPU VPS is enough for prototyping: Hostinger offers them starting at a few euros per month, enough to set up a VoiceStudio test bench before making a decision. But let's be clear-eyed: reaching Turbo's 150 ms in self-hosted mode requires infrastructure expertise that few teams have in-house.
Pricing: ElevenLabs' premium put to the test against real-world cost
At $80 per million characters, Eleven v4 is the most expensive model of the leading trio — nearly 5x Gemini 3.8 Flash TTS. This premium holds up for some use cases, but not all:
| Model | Price / 1M characters (September 2026) | Gap vs Eleven v4 |
|---|---|---|
| Eleven v4 | $80 (verify on elevenlabs.io) | — |
| Cartesia Sonic 3.6 | $49 (verify on cartesia.ai) | −39% |
| Gemini 3.8 Flash TTS | $16.49 (verify on ai.google.dev) | −79% |
Hourly cost: the only metric that really matters
One million characters corresponds roughly to 20 hours of spoken voice (estimated at 14 characters/second, to be recalculated based on your actual throughput). Eleven v4 therefore works out to about $4 per hour of voice, versus ~$2.45 for Sonic 3.6 and ~$0.80 for Gemini 3.8 Flash TTS.
Two factors narrow the gap, however. First, pronunciation robustness: fewer regenerations, hence a lower real cost — each billed retake doubles the price of the segment. Then, generation throughput (73.4 characters/s), which reduces machine time on batch pipelines.
My verdict by use case: for an audiobook or a brand voice, the $80 is justified — the quality of interpretation saves retakes that cost more in human time. For a high-volume agent loop, the gap multiplies with volume, and Gemini 3.8 Flash TTS or Sonic 3.6 become the rational choice.
Google, for its part, takes the cheap real-time logic to its conclusion with Gemini 3.8 Live, its conversational voice at $1.38 per hour — a positioning that ElevenLabs does not counter head-on, instead fully embracing the quality premium.
On the commercial side, the launch comes with two weeks of free generations for Creator+ plans in ElevenCreative, and monthly credits multiplied up to 2x. Enough to test v4 on your own corpus before pulling out the credit card.
Two tiers, one strategy: the most clear-eyed bet in the market
The two-tier strategy — a quality model, a latency model — is, in my view, the most pragmatic choice in today's voice market. And I bet it becomes the norm by 2027.
The market has split into two populations with incompatible needs: studios want emotion and finesse, agent developers want milliseconds. A single model claiming to serve both is bound to fail on one front — that's exactly what the competition shows: Cartesia bets everything on real-time, Google already separates Flash TTS (cheap) and Live (conversational).
The appeal of ElevenLabs' approach is the shared foundation: the same voice library, the same cloning pipeline, the same expressive tags across both tiers. A studio that produces a brand voice with v4 can deploy it as an agent with Turbo without redoing any upstream work.
The risk is real: maintaining two models means accepting a possible quality divergence between the two lines. But the co-optimization of Turbo with ElevenAgents shows that latency was treated as a design constraint, not as a mere deployment parameter.
One last structural advantage: ElevenLabs remains modular — a TTS layer that plugs into the LLM of your choice. Conversely, Gemini 3.8 Live bundles voice and reasoning into a single service. For a team that has already settled on its LLM stack (GPT-5.5, Gemini 3.1 Pro...), modularity weighs heavily in the decision.
Integration: a one-parameter switch, not a rewrite
On the integration side, switching between the two models is done with a single parameter — eleven_v4 or eleven_v4_turbo as the model_id — across all three ElevenLabs surfaces: ElevenCreative for creation, ElevenAgents for agents, ElevenAPI for development. The official model documentation details the IDs and capabilities of each.
No pipeline rewrite, no additional endpoint. It's a detail that matters: you can route intelligently — Turbo for live calls, v4 for content production — without maintaining two separate integrations.
My migration recommendation: don't switch everything over at once. Duplicate your routes, send 10% of traffic to Turbo, and measure three metrics with your real users: effective time-to-first-speech, the interruption rate (barge-in), and the early-call abandonment rate. That's where the 150 ms convert into euros.
TestingCatalog confirms that Turbo specifically targets live calls and agent loops — the product positioning is consistent with the benchmarks. There's no point paying for real-time where a batch pipeline suffices.
Voice agents and avatars: what this changes for developers
With Turbo at 150 ms, speech synthesis is no longer the bottleneck for voice agents. The time budget gets redistributed to the three other links in the chain that determine perceived fluidity: speech recognition, LLM reasoning, and tool execution.
To build a complete stack, you'll therefore need a good transcription engine upstream — our comparison of the best speech recognition AI reviews the options on the market. On the brain side, an agentic LLM like GPT-5.5 or Gemini 3.1 Pro handles the reasoning while Turbo generates audio in a continuous stream.
The avatar front, on the other hand, is not where ElevenLabs plays. Google added real-time visual presence to Gemini 3.8 Live with Live Avatar: voice and animated face in the same conversational loop. For a complete talking avatar, Google has a head start; for a purely voice-based agent — telephony, support — Turbo is the current latency benchmark.
More broadly, the race to real-time extends beyond voice: AlayaWorld, the first open source world model to generate playable worlds in real time beyond 60 seconds, shows that all of AI generation is converging toward instant, continuous outputs. Turbo is not an isolated product: it's a market symptom.
My overall take: synthetic voice is becoming technically commoditized, and differentiation is moving up the stack — orchestration, tools, conversational memory. ElevenLabs understood this by opening up v4 and Turbo across its three surfaces from day one.
❌ Common Mistakes
Mistake 1: using v4 for a real-time agent
Eleven v4 is optimized for quality, not latency: putting it in a conversational loop means paying the premium without getting real-time benefits. Solution: eleven_v4_turbo for anything live, v4 for content production. The two-tier strategy exists precisely for that.
Mistake 2: comparing prices without factoring in regenerations
TTS pipelines regenerate — retakes, botched punctuation, emotions that need redoing. At $80 versus $16.49 per million, a 20% regeneration rate widens the gap even further. Solution: budget with a retry coefficient measured on your actual corpus, never a theoretical one.
Mistake 3: cloning a voice without rights or consent
Cloning a voice from 10 seconds of audio is now trivial — and legally sensitive: voice rights apply, and platforms require the speaker's consent. Solution: systematic written authorization, and Professional Voice Clone for supervised professional use.
Mistake 4: taking the leaderboard's word for it
An Elo of 1,319 in a voting arena guarantees nothing for your use case — Controlled Voice proves it, with Alibaba outperforming ElevenLabs. Solution: test with your own scripts, voices, and listeners before migrating. Benchmarks are a filter, not a verdict.
❓ Frequently Asked Questions
What's the difference between Eleven v4 and Eleven v4 Turbo?
v4 is the most expressive model, designed for studio-quality output; Turbo is its real-time variant with ~100 ms median latency, co-optimized with ElevenAgents. Turbo retains expressive controls and the Professional Voice Clone during calls. Same API, switch via the model_id parameter.
How much does Eleven v4 cost?
$80 per million characters (September 2026, check elevenlabs.io). Turbo's pricing was not detailed at launch. Launch offer: free generations for two weeks and up to 2x monthly credits for Creator+ plans.
Does Eleven v4 handle French well?
Yes. The model covers 90+ languages and posts a record 91.7% pronunciation robustness score (versus 85.6% for v3). It interprets acting directions — laughter, sighs, accents, anger — and can make a voice recorded in French speak another language with a native accent.
Should you leave ElevenLabs for VoiceStudio or Gemini 3.8 Flash TTS?
It depends on your dominant constraint: data sovereignty or very large volumes → self-hosted VoiceStudio; budget above all → Gemini 3.8 Flash TTS at $16.49/1M; quality and real-time agents → ElevenLabs. In most cases, no migration is warranted — test before deciding.
What latency is needed for a natural voice conversation?
Research on turn-taking places the human pause at around 200 ms. Turbo, at 150 ms time-to-first-speech, gets back under that threshold on the synthesis side and leaves headroom for recognition and the LLM. Beyond 300 ms cumulative, the conversation feels artificial.
✅ Conclusion
Eleven v4 consolidates ElevenLabs' quality dominance, and Turbo pushes speech synthesis below the human pause threshold — if you're building voice agents, now is the time to redo your benchmarks, and the launch offer (free generations for two weeks, 2x credits for Creator+ plans) is the perfect excuse to test Eleven v4 right now.