ElevenLabs launched two new speech models on Monday, called ElevenLabs v4 and v4 Turbo. The company said they offer more expression control, lower latency for voice agents and support for more than 90 languages.
The previous version supported 70 languages. ElevenLabs said it observed the biggest quality jump in Japanese, Brazilian Portuguese, Mandarin and Cantonese.
For the v4 generation, ElevenLabs is adopting a new architecture that it says allows for better control and faster cloning. The company said users will be able to clone a voice with just 10 seconds of audio.
On the creative side, the model handles voice identity better over longer chunks of text and keeps the context of the text in mind while reading it aloud to change expressions. ElevenLabs introduced inline tags to define expression with v3 and is expanding those tags in v4, letting users stack multiple tags and having the model follow the sequence.
ElevenLabs said the new model is suited for voice agents, as it has lower latency to allow for more fluid conversation. The v4 can start generating audio as soon as the LLM behind it starts generating answers, and the company said it can handle confrontations, escalations and holds differently for better issue resolution.
The company said more than 55% of its business comes from large companies. ElevenLabs raised $500 million earlier this year in a round led by Sequoia that valued the company at $11 billion. Its annualized revenue run rate has climbed from roughly $330 million at the start of the year to over $600 million, and headcount has reached over 800.