ElevenLabs
ElevenLabs v4 targets performed dialogue for AI video
The new model stacks direction tags, follows the scene it is in and clones a voice from ten seconds of audio, and its Turbo variant answers in about 150 milliseconds. Dialogue was the last unperformed layer of a generated scene.