ElevenLabs releases v4 speech models with 90 languages and lower latency
ElevenLabs launched v4 and v4 Turbo speech models with more expression control, faster voice cloning, lower latency for voice agents and support for more than 90 languages, as its enterprise business and revenue grow.
The company released its v3 model last year and teased the new model at an event in Warsaw earlier this year. For the v4 generation, ElevenLabs is adopting a new architecture that allows for better control and faster cloning. The company said users will be able to clone a voice with just 10 seconds of audio.
On the creative side, the model handles voice identity better over longer chunks of text. It also keeps the context of the text in mind while reading it aloud to change expressions. ElevenLabs introduced inline tags to define expression with v3 and is expanding those tags in v4, letting users stack multiple tags and having the model follow the sequence.
The previous version supported 70 languages, and ElevenLabs has worked to get that number up to more than 90 languages with the new version. The company said it observed the biggest quality jump in Japanese, Brazilian Portuguese, Mandarin and Cantonese.
ElevenLabs has scaled its enterprise calling business rapidly over the last year, with more than 55% of its business coming from large companies. The company said the new model is suited for voice agents because the new version has lower latency to allow for more fluid conversation. The v4 can also start generating audio as soon as the large language model behind it starts generating answers. The model can handle confrontations, escalations and holds differently for better issue resolution.
Competition in speech models has ramped up as startups such as Cartesia, Deepgram, Fish Audio, Boson and WellSaid Labs have created expressive speech models. Large companies including Google and OpenAI have also improved their voice models.
ElevenLabs raised $500 million earlier this year from Sequoia in a round led by Sequoia that valued the company at $11 billion. There are already rumors of a followup fundraising round that would value the company at $22 billion. ElevenLabs’ annualized revenue run rate has climbed from roughly $330 million at the start of the year to over $600 million. The company has aggressively hired personnel in markets such as India, Europe and Brazil, and its headcount has reached over 800.
In a recent interview with TechCrunch, the company’s co-founder and CEO Mati Staniszewski said the company is aiming for an IPO in the next years, but he did not commit to a timeline.
Editor's Summary ElevenLabs launched v4 and v4 Turbo, adding expression controls, faster voice cloning, lower latency and support for more than 90 languages. The release comes as the company reports enterprise growth, rising revenue and IPO ambitions, with competition from other speech model developers and large technology companies.