ElevenLabs v4 targets performed dialogue for AI video
The new model stacks direction tags, follows the scene it is in and clones a voice from ten seconds of audio, and its Turbo variant answers in about 150 milliseconds. Dialogue was the last unperformed layer of a generated scene.

Key takeaways
- ElevenLabs launched Eleven v4 and Eleven v4 Turbo on September 28, 2026, available in ElevenAgents, ElevenCreative and the API, with access on the free tier.
- Version 4 clones a voice from 10 seconds of audio, supports 90+ languages against 70 in v3, and expands the inline audio tags that direct delivery, with SSML controls such as the break tag disabled in their place.
- Turbo is built for real time, with a median inference latency of about 100 milliseconds and about 150 milliseconds to the first speech, which is inside the pause between two people talking.
- ElevenLabs says v4 ranks first on Artificial Analysis's provider voice arena, above Cartesia's Sonic 3.6 and Google's Gemini 3.8 Flash TTS, and that listeners preferred it in about 75 percent of blind head-to-head tests.
ElevenLabs shipped two speech models on September 28, 2026: Eleven v4 and Eleven v4 Turbo, both available immediately in ElevenAgents, ElevenCreative and through the API, and both included on a free account tier. The company calls v4 its most emotive model yet, which is marketing language, and this time the technical description behind it says something specific.
Version 3 could read a script well. Version 4 is built to read it the way an actor does. That is the claim, and for anyone assembling a film out of generated scenes it is aimed at the weakest part of the pipeline.
What actually changed

The company's own chart for v4: it reports wins in 65 to 81 percent of blind head-to-head preference tests against competing speech models. Study by ElevenLabs. Image: ElevenLabs.
Direction in the script, and it stacks. ElevenLabs introduced inline tags for delivery with v3 and has expanded them. A line can now carry several, in sequence: an emotion, a pace, an accent, a sound. The examples on the product page and in the launch post run from [whispers] and [nervous laugh] to [door slams] and [gong sounds], and the company says v4 follows a sequence of tags more reliably than v3 did, sound effects included. The trade is that SSML is out: the documentation says controls such as the break tag are disabled in v4, with natural-language tags as the replacement.
Context, not isolated lines. Multi-speaker dialogue is generated as a scene rather than as separate reads stitched together, so a speaker responds to what the previous line just said. Anyone who has cut a two-hander out of separate text-to-speech passes knows why that matters: the performances used to arrive from different rooms.
Long-form without the seams. Context stitching is meant to keep pacing and delivery steady across a script of any length, which is aimed at audiobooks and, more interestingly for video, at narration that runs for forty minutes without drifting.
Regenerating a line does not move the voice. Speaker stability across repeated generations is the boring-sounding feature that decides whether a generated character can hold a scene, and ElevenLabs says it holds across dialogue and narration.
Cloning from ten seconds. Version 4 can clone a voice from about 10 seconds of audio, and professional voice clones are back, which v3 did not support, now carrying the model's full emotional range in every language the clone speaks.
Languages: 70 to 90-plus. The company says the largest quality jumps are in Japanese, Brazilian Portuguese, Mandarin and Cantonese. The character limit for a single v4 request is 10,000, against 5,000 for v3, per the model documentation.
The realtime half is the part video should watch
Turbo is the variant aimed at conversation rather than production. ElevenLabs gives a median inference latency of roughly 100 milliseconds, with about 150 milliseconds to the first speech, which it points out is inside the average pause between two people talking. It was tuned together with ElevenAgents rather than bolted onto it.
That number matters beyond call centres. The interactive side of AI video, the world models and real-time avatar systems we have covered this month, all assume a voice that answers in the time a person would. A model that generates the first syllable in 150 milliseconds removes the last obvious reason those characters feel like machines, and it does it over an API rather than inside a game engine.
Why this is a video story
An AI-generated shot has been cheap for a while. What has stayed expensive and slow is everything that turns shots into a scene: continuity, structure, and dialogue that lands. Voice was the layer where a viewer's tolerance snapped, because the ear is unforgiving about performance in a way the eye is not.
Four of these changes point straight at that:
- Performed dialogue in one pass makes a two-hander possible without a cast, which is the difference between a narrated short and a scene.
- Consistent identity across regenerations is the audio equivalent of character consistency, the problem that eats generated films at the edit stage.
- Ninety-plus languages with professional clones means a finished film localises into most of the world's markets with the original performance intact, which is the workflow our own dubbing test covered and which used to require a cast per language.
- 150 milliseconds to first speech makes a character answerable in real time, which is what interactive video needs and what a rendered clip never required.
The honest caveat is ours: SLOP TV has not tested v4. We have no production pass with it, no same-script comparison against the voice our own films use now, and no measurement of our own. Everything above is the company's description of its own model plus the third-party notes that exist, and a first look is not a review.
The leaderboard claim, read carefully
The headline number in the launch post is that Artificial Analysis ranks v4 first on its provider voice arena, and the company says listeners preferred it in about 75 percent of blind head-to-head tests. The arena figure is corroborated from the other side: Artificial Analysis's own post on September 28 puts Eleven v4 first on the Provider Voice TTS Arena and on Pronunciation Robustness, and second on Controlled Voice, ahead of Cartesia's Sonic 3.6 and Google's Gemini 3.8 Flash TTS.
Three things that does not tell you. The arena is a human-preference result, so it measures how speech sounds rather than how a model behaves across a forty-minute script. The blind-test percentage is the maker's own study. And a leaderboard win over a Google model is a fact about relative quality on that bench, not a statement about who ships the best film.
What to watch next
- Pricing and credit cost for long-form: the launch post does not answer what a feature-length read costs, and 10,000-character limits mean a script is still assembled in passes.
- Whether the tag system survives a long script. Stacking works in the demos; the question is the fifth pass of a dialogue scene, not the first.
- Independent listens. The arena tells you a preference; it does not tell you where the model breaks. That is what a same-script test against two rivals would show, and nobody has published one yet.
- The dubbing pipeline. If a professional clone can carry full emotional range across 90 languages, the economics of localising a finished video change for everyone who makes one.
If you make video with generated voices, take one scene you have already cut and re-read both sides of the dialogue as a single tagged block. That is the change, and it is testable in an afternoon.
Sources
- elevenlabs.io - the product page: the new architecture, tag stacking, context stitching, professional voice clones, Turbo latency
- elevenlabs.io - the launch post: Artificial Analysis first place, the blind preference numbers, tag examples
- elevenlabs.io - the model list: character limits, languages, the v4 and Turbo descriptions
- techcrunch.com - the launch report: the 10-second clone, the language jump and where the quality improved
- unite.ai - the tag and SSML details, availability on the free tier
- 247wallst.com - the Artificial Analysis leaderboard post, ranked ahead of Sonic 3.6 and Gemini 3.8 Flash TTS
- sloptvnews.com - our own how-to on dubbing a finished video, for the workflow context