NEWS

ElevenLabs launches v4 with more expression control and support for 90 languages

The new v4 and v4 Turbo models reduce latency for voice agents, allow voice cloning with 10 seconds of audio, and bring the largest quality leap in Brazilian Portuguese the company has ever recorded.

ElevenLabs announced on Monday (September 28) two new speech synthesis models, v4 and v4 Turbo, which expand control over speech expression, reduce latency for voice agents, and raise the number of supported languages from 70 to 90. The launch was reported by TechCrunch and marks the direct successor to v3, launched last year and unveiled at a company event in Warsaw.

For those building products with speech synthesis, the most relevant change isn't cosmetic: it's architectural. ElevenLabs adopted a new model architecture for v4 that, according to the company, allows cloning a voice with just 10 seconds of audio and better preserves vocal identity across long texts, something earlier models tended to lose over lengthy paragraphs.

Expression tags gain chaining

v3 introduced inline tags to control expression in synthesized speech, something like marking a piece of text to be read in a certain tone. v4 expands this feature: it's now possible to stack multiple tags, and the model follows their sequence, while also keeping the text's context in mind as it reads, adjusting expression based on what comes before and after. In practice, this should significantly reduce the work for those who currently break voice scripts into smaller chunks just to manually control intonation, a recurring problem in automated narration and synthetic dubbing pipelines.

Brazilian Portuguese among the biggest quality leaps

ElevenLabs stated it observed the biggest quality leap in four languages: Japanese, Brazilian Portuguese, Mandarin, and Cantonese. For those developing voice products for the Brazilian market, this matters more than the round number of 90 languages: it's precisely in the languages historically most difficult for speech models (tones, prosody, regional variation) that the company says it advanced the most in this generation. This is relevant because, so far, most public TTS (text-to-speech) benchmarks prioritize English and a handful of European languages, leaving Brazilian Portuguese as fertile ground, but one rarely tested publicly by direct competitors.

Lower latency designed for voice agents

ElevenLabs has invested heavily in corporate customer service via voice calls: more than 55% of the company's revenue already comes from large enterprise accounts, according to TechCrunch. v4 was designed with this in mind. Lower latency allows for more fluid conversations, and the model can start generating audio the moment the LLM behind the agent begins generating its response, instead of waiting for the full text to be ready before synthesizing the voice. This is a concrete technical difference for those building voice agent pipelines: it eliminates a sequential waiting step (LLM finishes -> TTS starts) and enables parallel streaming instead, reducing the time to the first audible audio.

The company also says v4 handles confrontation, escalation, and hold situations on calls differently, common scenarios in automated call centers, where tone of voice needs to change depending on whether the customer is upset, waiting to be transferred, or in the process of resolving an issue.

A market that keeps growing

The launch of v4 comes at a moment of accelerated growth for ElevenLabs. The company raised $500 million in a round led by Sequoia earlier this year, valuing it at $11 billion, and there are already rumors of a new round that would value it at $22 billion. Annualized revenue (ARR) jumped from about $330 million at the start of the year to more than $600 million, and headcount surpassed 800 people, with aggressive hiring in markets such as India, Europe, and Brazil. In a recent interview with TechCrunch, CEO and co-founder Mati Staniszewski said the company intends to go public "in the coming years," without committing to a date.

Competition in the expressive speech model space has grown alongside it: startups like Cartesia, Deepgram, Fish Audio, Boson, and WellSaid Labs have been releasing their own models, while Google and OpenAI continue improving their native voice models. This means that anyone choosing a TTS API today has more options than a year ago, and the decision increasingly depends on specific technical criteria (real production latency, Portuguese quality, cost per audio token) rather than picking "the one vendor that does it all."

What remains open

The announcement doesn't include latency numbers in milliseconds, nor published comparative benchmarks against Cartesia, Deepgram, or Google's and OpenAI's voice models. There is also, as of the source's publication, no public technical documentation detailing usage limits for the new stacked expression tags or specific pricing for v4 and v4 Turbo for Brazilian developers. Those already using the ElevenLabs API in production will need to test in practice how much the quality leap in Brazilian Portuguese translates into less manual adjustment of pronunciation and intonation, something that only shows up when running one's own production scripts against the new model.

Translated from the Brazilian Portuguese original · Read the original

Read also