ElevenLabs, a company that makes AI voices, released Eleven v4 on Sept. 28. It is a model that turns written text into speech, and the company says it is built to read a script the way a voice actor would: it works out who is speaking, what just happened and how each line should land, so a line can come out tender, urgent or funny. ElevenLabs' launch video opens by saying that everything in it was generated from the prompts shown, with no edits.
Users can also direct the performance by writing instructions in square brackets, such as [laughs], [whispering] or [light rain], and the model is meant to add the laugh, the whisper or the sound. It can voice a scene with several speakers, too, and ElevenLabs says each one now reacts to what was just said instead of sounding like separate lines stitched together. The company says the model follows these directions better than its older ones, but also that its handling of them is "not perfect yet."
The model can copy a voice as well. ElevenLabs says a quick copy can now be made from just 10 seconds of audio, and that a voice recorded in one language can speak any of the more than 90 languages the model supports, with the accent of a native speaker. The company says every copied voice needs verified consent from the person it belongs to, and that it can detect audio made with the model as AI-generated.
Eleven v4 is available now in ElevenLabs' apps and for developers to build into their own products, and ElevenLabs says it costs the same as the company's other voices, which puts it on the free plan too. A faster version, Eleven v4 Turbo, is meant for voice assistants that talk with people in real time, such as ones that answer callers.
On a public leaderboard run by Artificial Analysis, where listeners pick between two voices without knowing which company made them, the two new models hold the top two spots, close enough that the ranking treats them as roughly tied. That ranking measures which voice people find more natural, not whether the model reads every line right.
ElevenLabs says the accent switching is always on for now, so a copied voice speaking a new language loses its original accent unless written directions can coax it back, which the company says may not work. Letting users turn the switching off is still a research project with no timeline, and ElevenLabs says it is still training the model, so how it behaves may shift over time.
