ElevenLabs launches Eleven v4 with tighter voice control
ElevenLabs releases Eleven v4, which follows laughter and whisper cues more accurately and keeps voices stable across audiobooks. A Turbo variant responds in 150 milliseconds.

Updated
Why it matters
- ElevenLabs released Eleven v4, a speech model that follows cues for laughter and whispering more accurately and maintains voice consistency across long productions like audiobooks.
- The Turbo variant of Eleven v4 starts speaking in 150 milliseconds and is built for real-time voice agents.
- On Artificial Analysis' Voice Arena leaderboard, Eleven v4 ranks ahead of Cartesia and Google's Gemini.
ElevenLabs has released Eleven v4, a speech model that follows cues for laughter and whispering more accurately and keeps voices consistent across long productions such as audiobooks. The release targets the two weaknesses that have historically separated synthetic speech from human narration: emotional expressiveness and long-form stability.
The company also shipped a Turbo variant of the model, which starts speaking in 150 milliseconds. That latency figure matters for a specific and fast-growing market: real-time voice agents. Voice interfaces that pause for a second before responding feel broken to users, and sub-200-millisecond response times put Eleven v4 Turbo in the range needed for natural-sounding conversations between humans and AI agents.
The new model arrives with an external validation. On Artificial Analysis' Voice Arena leaderboard, Eleven v4 ranks ahead of Cartesia and Google's Gemini. That comparison places ElevenLabs ahead of two of its most serious competitors in the speech generation market — Cartesia, a startup focused on low-latency voice models, and Google, which has been folding speech capabilities into its Gemini model family.
Why expressiveness and consistency are the battleground
The improvements ElevenLabs highlights in v4 map directly onto the two main commercial use cases for AI speech. Expressiveness — the ability to laugh, whisper, and shift tone on cue — determines whether a synthetic voice can carry entertainment and narration content without sounding flat. Consistency across long productions determines whether publishers can use AI voices for audiobooks, where a voice that drifts or degrades over hours of audio is unusable.
These have been the hardest problems in text-to-speech. Short demo clips of AI voices have sounded convincing for years; maintaining character over a ten-hour audiobook, or rendering a whispered line exactly where a script demands it, is where models have tended to fail. By framing v4's release around these specific capabilities, ElevenLabs is positioning the model for production work rather than demos.
The Turbo variant's 150-millisecond response time addresses a different pressure point. Voice agents — AI systems that hold spoken conversations with customers or users — have become one of the most competitive segments of the AI market, and latency is the gating factor. A model that begins speaking in 150 milliseconds keeps conversation rhythm close to human turn-taking.
A crowded field
The leaderboard placement underscores how contested the speech generation market has become. ElevenLabs built its business on high-quality synthetic voices, but it now faces competition on two fronts: specialized startups like Cartesia that optimize for speed, and hyperscalers like Google that bundle speech into broader multimodal model platforms. Ranking ahead of both on Artificial Analysis' independent benchmark gives ElevenLabs a concrete data point in that fight.
The launch also signals where the company sees demand heading. Audiobook production and real-time agents represent two of the largest commercial opportunities for synthetic speech — one in publishing and media, the other in customer service and enterprise automation. A single model family serving both, with a standard version tuned for quality and a Turbo version tuned for speed, suggests ElevenLabs intends to cover the full spectrum rather than cede either segment.
With v4 now released and benchmarked ahead of Cartesia and Gemini on Voice Arena, the question shifts to how quickly rivals respond — and whether Google and others can close the gap in expressiveness and long-form consistency that ElevenLabs has targeted with this release.
Original: elevenlabs.io
More from Elena Vasquez
Show full bio
Market editor covering media and advertising at AI In Context.
143 articles
Related articles
- Google Ships Upgraded Gemini 2.5 Flash Native Audio and Live Translation
- Google Ships Gemini 3.8 Live and Extended Thinking Models
- Google Launches Gemini 3.1 Flash TTS With Audio Tags and 70+ Languages
- Google Ships Gemini 3.8 TTS Models With Voice Cloning and Direction
- Google's Flash TTS models build AI voices from text descriptions