Skip to main content

Eleven v4: What ElevenLabs' New Voice Model Changes for AI Ad Voiceovers

ElevenLabs launched Eleven v4 and a low-latency Eleven v4 Turbo on September 28, 2026. Here is what audio tags, two-speaker scenes, 90+ languages and 10-second clones change for a 30-second ad read, what a take costs on fal, and a test an ad team can run in one afternoon.

Mauricio Valdivia

Mauricio Valdivia

·13 min

An ad producer in studio headphones listens to voiceover takes beside a marked-up printed script, a laptop showing three audio waveforms and an amber bottle

Your script is fine. The read is flat.

A cold-brew brand has a 30-second script that already won once with a human creator. Nothing is wrong with the words. The AI voiceover version reads it like a terms-of-service page: every sentence lands at the same speed, the joke gets no laugh, and the product name comes out slightly wrong.

On September 28, 2026 ElevenLabs launched Eleven v4, which its announcement calls "our most emotive text-to-speech model yet," together with a low-latency variant named Eleven v4 Turbo. The company had previewed v4 at an event in Warsaw in June. September 28 is the launch, and ElevenLabs says both models are available now in its own apps and through its API.

For people who make UGC-style ads, four changes in the announcement matter: you can direct delivery with inline tags and plain-language notes, two speakers can share one scene, one voice can speak more than 90 languages with a native accent, and an instant clone now starts from 10 seconds of audio. This post reads the launch for a 30-second ad read, prices a take on fal, and ends with a test an ad team can run in one afternoon. Novoads does not offer Eleven v4. This is a read of someone else's model, written for the people who will direct it.

What ElevenLabs launched on September 28

Two models, one architecture, split by the job they do. Before the ad-specific parts, here is what shipped and which numbers are whose.

Two models: one for produced audio, one for agents

ElevenLabs draws the line itself. Its product page says "Eleven v4 is tuned for produced content where quality matters most," and describes Eleven v4 Turbo as "a low-latency variant at ~100 ms median inference latency, designed for voice agents and real-time use." The same page says both models share the same expressive range and both support Professional Voice Clones.

Our reading: an ad voiceover is produced content. Nobody is waiting on the other end of a phone line while your 30-second read renders, so Turbo's speed buys nothing a viewer can hear. The ad model is v4. Turbo matters to you only if your funnel includes a voice agent, a live sales line or an interactive demo.

What ElevenLabs says changed since v3

The v4 FAQ answers what is new compared with Eleven v3 in a single paragraph. According to ElevenLabs, Eleven v4:

  • is built on a new architecture with higher audio quality and a wider emotional range
  • keeps speaker identity stable across regenerations
  • supports Professional Voice Clones again, after v3 did not
  • follows audio tags more reliably
  • has a lower time to first byte

The launch post adds two details that matter for production work. Support for International Phonetic Alphabet (IPA) phonemes "has also been significantly improved," and request stitching, which chains generations together for longer content, is described as significantly more reliable. TechCrunch, covering the launch the same day, reported that ElevenLabs "launched two new speech models on Monday," which matches the September 28 date on the announcement.

The rank, the blind test and the latency are ElevenLabs' numbers

Three figures from the launch will be repeated everywhere this week. All three are ElevenLabs' claims about its own model, and each measures something specific:

  • Rank. ElevenLabs says v4 is ranked #1 by Artificial Analysis, citing the Provider Voice Arena Leaderboard for September 2026.
  • Blind test. ElevenLabs says v4 was preferred by ~75% of listeners in blind head-to-head tests. Its footnote names the comparison set as Cartesia Sonic 3.6, Inworld TTS-2, Google Gemini 3.8 Flash-Lite TTS and Google Gemini 3.8 Flash TTS in September 2026, says graders judged which line was more expressive and which sounded more natural, and says ties counted as half.
  • Latency, two metrics. For Turbo, ElevenLabs quotes a median inference latency of ~100 ms. Separately, it quotes a median time to first speech of ~150 ms, which its footnote defines as the median time from request to audible speech, measured with Turbo over WebSocket streaming.

Keep the two latency numbers apart. Inference latency is the model's own processing time. Time to first speech, per ElevenLabs' footnote, runs from request to audible speech with network latency measured and removed, so a real caller waits for that figure plus the network. Neither one changes a pre-rendered ad file. The field also moves fast: Google's Gemini 3.8 Flash TTS was announced on September 23, five days before this launch, and already sits in ElevenLabs' comparison set, while in July the voice-model headline belonged to Alibaba's Qwen text-to-speech release. A rank is a reason to audition a model, not a reason to switch your ads to it.

A real UGC creator filming herself on a phone
Novoads · UGC video ads with AI, ready in minutes.
Try now

Directing a 30-second ad read

The biggest change for ad work is control. A flat read has always been fixable by re-recording a human; with a synthetic voice, the fix has to live in the script.

Plain-language direction and inline tags

ElevenLabs says users "can describe how a line should be delivered in natural language," and can add inline tags like [laughs], [said angrily in French accent], [light rain] or [phone buzzing] to set emotion, phrasing and even sound effects. The announcement says "Eleven v4 follows these audio tags and direction prompts more accurately than prior models," and the product page says v4 "follows tag sequences more reliably than v3, sound effects included."

Two details from the v4 FAQ matter if you are porting older scripts:

  • Pauses are tags now. The FAQ says SSML break tags are disabled in v4, so you write [pause] or [long pause] where you want a break.
  • Length is not the constraint. A single generation supports up to 10,000 characters. A 30-second ad script sits far below that ceiling.

The announcement says ElevenLabs' developer documentation covers the full tag syntax and how to apply these tags through the API.

A worked example: one cold-brew ad, three reads

Here is a UGC-style script for a cold-brew concentrate, marked up for v4. It is 48 spoken words, about 330 characters with its tags, so it runs short of a full 30-second read:

[warm] Okay, honest review. [pause] I was spending way too much on café cold brew. [sighs] Every single week.
[excited] Then I tried this concentrate. One pour, water, ice, done.
[laughs] My barista definitely misses me.
[pause] It tastes like the café version, and one bottle lasts me two weeks.
[whispers] Don't tell my barista.

Render it three ways and hold everything else fixed, same voice and same settings:

  1. Plain text, no tags. This is your baseline, roughly the flat read you started with.
  2. Tagged, as above. The tags carry the arc: a skeptical open, a small complaint, relief, a laugh, a quiet close.
  3. Plain text plus a one-line delivery note, for example "a friend recommending something over coffee, relaxed and a little amused, picking up speed on the product line."

What you are listening for: does the laugh sound like a laugh, or like someone reading the word? Does the whispered close still carry the joke clearly? A read that nails the emotion but blurs the product is a worse ad than a flat one. The hook is the first line either way, so the principles in our guide to writing ad hooks apply to the delivery as much as to the words.

Brand names, IPA and pronunciation dictionaries

A mispronounced product name is the one error every viewer notices. ElevenLabs says IPA phoneme support "has also been significantly improved, so custom pronunciations behave more reliably," and its FAQ says "Pronunciation dictionaries still define how names and technical terms are spoken." For an ad team the practical rule is simple: define the brand name once, then check it in every take, including every regeneration. A dictionary entry you set in week one and never re-check is how a wrong vowel ends up in fifty variants.

Two speakers in one scene: skit and podcast-style ads

Single-voice reads are most of UGC. The formats that feel most organic, though, are often conversations, and that is where the launch makes its second claim.

What "understands the whole scene" means for dialogue

ElevenLabs says that because v4 understands the context of a whole scene, "it generates natural dialogue where speakers respond to what's just been said, rather than stitching together isolated lines." The launch post also says more natural multi-speaker dynamics make conversations feel responsive, "rather than like separate lines assembled together."

The common workaround for a two-voice ad has been to render each speaker's lines separately and cut them together in the edit. That is where the dead air between lines comes from, and why the second speaker so often sounds like they did not hear the first. If the scene-level claim holds on your script, the edit gets shorter and the timing stops giving the ad away.

Where a two-voice format earns its place

Four ad formats depend on two voices sounding like they are in the same room:

  • Skit, problem to solution. One person complains; the other hands over the product.
  • Podcast clip. A host and a guest, with the product as the thing the guest keeps bringing up.
  • Street interview. A quick question and an answer that sounds unscripted.
  • Support-call roleplay. ElevenLabs' own cloning demo is a support agent sorting out a double charge, which is a ready-made problem-and-fix structure.

What to check in a dialogue take

Before a dialogue take goes anywhere near an ad account, listen for four things:

  • Turn-taking. Does speaker B react to speaker A's line, or simply start talking?
  • Identity. ElevenLabs says v4 "preserves speaker identity more reliably across generations, dialogue, narration, and regenerated lines." Regenerate the scene twice and check that both voices are still themselves.
  • Tag placement. Laughs and pauses should land where the script puts them, not a line late.
  • Length. A skit that runs 42 seconds is a different ad from the 30-second one you briefed.

The voice is half of the continuity problem. The face is the other half, which we cover in keeping one AI actor consistent across ad variations.

Several UGC creators filming product variations to camera
Novoads · UGC video ads with AI, ready in minutes.
Try now

One brand voice in 90+ languages

Localization is where the announcement speaks most directly to advertisers, and where a test with real listeners matters most.

A native accent that keeps the voice

"Both Eleven v4 and Eleven v4 Turbo support more than 90 languages," ElevenLabs says, and a voice recorded in one language now speaks any other fluently, "adopting the accent of a native speaker while retaining the identity of the original." The product page puts it as a promise to point the models at Japanese, Spanish or Portuguese text and hear it spoken "fluently, with a native accent."

The announcement makes the ad case itself. For dubbing and localization, it says, the voice of the brand you have chosen, whether a celebrity you work with or a tone that fits your company, "will sound great in any language supported by Eleven v4." That is a vendor's promise. The next two sections are how to check it.

Accent drift in a 30-second read

ElevenLabs says accent adherence is noticeably stronger than before, so "the voice no longer drifts back toward its source accent over the course of a generation." In an ad, the last five seconds carry the offer and the call to action. A voice that slides back to its home accent there fails exactly where it costs the most, so the end of the read is where to listen hardest.

What to audition before you localize

Language is not market. Before a localized read ships, check it against the market it is for:

  • Variety. Mexican and Castilian Spanish, or Brazilian and European Portuguese, are different audiences. A native speaker of the target market scores the accent, not a native speaker of the language in general.
  • The brand name. Decide whether the voice says it the local way or the brand's way, and check it stays consistent.
  • Numbers and units. Prices, sizes and dates should be read the way that market reads them.
  • Idiom. A literal translation read with a perfect accent still sounds like a translation. Rewrite the script per market, then voice it.

If you are weighing other voice models for this job, our round-up of open-source ElevenLabs alternatives for ad voiceovers covers the self-hosted route and its trade-offs.

Cloning is where a brand voice or a creator's licensed voice lives, so the cloning changes carry the most operational weight for teams that already use one.

Instant clones from 10 seconds, and Professional clones return

ElevenLabs says "Instant Voice Clones can now capture voices with high fidelity using just 10 seconds of audio," and that "Eleven v4 also adds support for Professional Voice Clones (PVC), for the highest-fidelity cloning use cases." Its FAQ goes further and says Instant Voice Clones in v4 "now outperform the Professional Voice Clones of Multilingual v2." That last line compares ElevenLabs' own models with each other. Hear it on your voice before you rely on it.

Existing clones need retraining

The line most teams will miss sits in the FAQ: "For PVCs and IVCs created prior to Eleven v4's launch, to work effectively you will need to retrain them with Eleven v4." If your brand voice, or a creator's voice you license, was cloned on an older model, plan for a retrain and a fresh approval from whoever signs off on that voice before any ad's audio moves to v4. The ElevenLabs voice library is not affected the same way: the FAQ says every voice in the library, which it counts at 17,500+, works with Eleven v4.

ElevenLabs says "Every clone requires verified consent from the voice's owner," and that generated audio is covered by AI Speech Classifier technology that can detect it as AI-generated. For a creator's voice, that turns licensing into a step the creator takes part in, not a file you hold. Google took a similar line with the consent recording behind Gemini 3.8 Flash TTS voice replication.

Detection is not disclosure. A classifier that can flag audio as synthetic does not answer an ad platform's own rules on AI content, which are set out in whether you have to label AI-generated ads.

Eleven v4 or v4 Turbo, and what a take costs

With the capabilities covered, the two decisions left are which model to call and what a round of takes will cost.

Use v4 when, use Turbo when

Use Eleven v4 when:

  • you are rendering a finished voiceover for a video ad, a UGC read or a dub
  • you are directing emotion line by line and regenerating until the read lands
  • the piece is long enough that stitching and speaker stability start to matter

Use Eleven v4 Turbo when:

  • the voice has to answer in real time, as in a sales or support agent or an interactive demo
  • you stream text in from a language model and need audio before the sentence ends, which the product page describes as pushing text "as your LLM generates it"
  • you generate at very high volume and want fal's lower per-character price, after checking it holds up on your own script

What a 30-second read costs on fal

As of September 29, 2026, fal hosts both models, and its model pages list these rates; they are fal's prices, not ElevenLabs'.

Eleven v4Eleven v4 Turbo
Built for, per ElevenLabsProduced contentVoice agents, real time
fal price per 1,000 characters$0.08$0.04
One 500-character takeAbout 4 centsAbout 2 cents
Fifty takesAbout $2About a dollar

The per-take rows are our arithmetic. A 30-second read runs about 70 to 80 spoken words, roughly 400 to 450 characters, and tags push a marked-up script toward 500. We assume tags count toward the character total, because they travel in the same text field as the script. On ElevenLabs' own plans the company says "Eleven v4 uses the same credit pricing as our other Text to Speech models."

The stance we would take from that table: the voice is the cheapest line item in the ad. Fifty takes cost about as much as a coffee. The expensive part is the listening, since fifty takes is an hour of someone's attention, and that is the argument for a fixed protocol instead of auditioning by feel.

Where each model runs

ElevenLabs' September 28, 2026 announcement says "Both models are available now in ElevenAgents, ElevenCreative, and via ElevenAPI." On fal, read on September 29, 2026, the Eleven v4 page lists audio tags, voice stability, similarity settings and IPA pronunciation as its delivery controls. Before you run the third read there, check that fal accepts a plain-language delivery note.

Real UGC creators talking to camera in a row of video cards
Novoads · UGC video ads with AI, ready in minutes.
Try now

A one-afternoon test for ad teams

A launch post is a demo. Your ad account is the test. Here is a protocol small enough to run before the end of the day.

Set up the test

  1. Pick one script that already runs, with a hook rate you know. If you need the formula, it is in our hook rate explainer.
  2. Keep the current audio as the control. Whatever voice the winning ad uses today is the bar.
  3. Render the three reads from the worked example on v4, with the same voice for all three.
  4. Render one localized version in your second-biggest market's language, with the same voice.
  5. Regenerate the best take five times and check the brand name in all five.

Score it like an ad, not a demo

  • Blind listen with three people from the audience. Ask one question: which of these sounds like a person talking to you?
  • Send the localized take to a native speaker of that market and ask about the accent and the offer line only.
  • Only then go live. Run the top two reads against the control as variants on the same visuals and equal spend, the same way you would test ten hook variations on one ad.

How to know it worked

The new read beats the control on hook rate and hold rate at equal spend, and the brand name survived every regeneration. If a read wins in the listening room but not in the ad account, the voice was not your bottleneck, and the next test belongs on the script or the first frame.

Where Novoads fits: the voice inside a UGC ad

Novoads does not offer Eleven v4, and nothing in this post describes a Novoads feature. In Novoads the voice is one layer of the video rather than a separate file: you write or paste a script, pick an AI actor, and the talking actor delivers it with an AI voice, lip-sync and captions in the same workflow, with voices in 31 languages and their regional accents.

The protocol above works the same way there. One script, several reads, the current winner as the control, and the ad account as the judge. You can run it in Novoads, and every plan is listed on the pricing page.

Eleven v4's real change for advertisers is not a leaderboard position. It moves delivery from something you accept to something you direct, and anything you can direct, you can test.

Frequently Asked Questions

What is Eleven v4?

Eleven v4 is a text-to-speech model ElevenLabs launched on September 28, 2026. Its announcement calls it 'our most emotive text-to-speech model yet' and launches it together with a low-latency variant, Eleven v4 Turbo. ElevenLabs had previewed v4 at an event in Warsaw in June; September 28 is the launch. The company's September 28, 2026 announcement says both models are available now in ElevenAgents, ElevenCreative and via ElevenAPI.

What is the difference between Eleven v4 and Eleven v4 Turbo?

ElevenLabs says Eleven v4 is tuned for produced content where quality matters most, and Eleven v4 Turbo is a low-latency variant designed for voice agents and real-time use. It quotes two separate speed figures for Turbo: a median inference latency of about 100 ms, and a median time to first speech of about 150 ms, which it defines as the time from request to audible speech. For a pre-rendered ad voiceover, v4 is the natural fit.

Can I direct the emotion of an ad voiceover in Eleven v4?

Yes, according to ElevenLabs. You can describe how a line should be delivered in natural language and add inline tags such as [laughs], [whispers] or [pause] inside the script, and ElevenLabs says v4 follows these tags and direction prompts more accurately than prior models. Its FAQ says SSML break tags are disabled in v4, so pauses are written as [pause] or [long pause].

How many languages does Eleven v4 support, and does the voice keep a native accent?

ElevenLabs says both Eleven v4 and Eleven v4 Turbo support more than 90 languages, and that a voice recorded in one language can speak another while adopting the accent of a native speaker and keeping the identity of the original. It also says the voice no longer drifts back toward its source accent over the course of a generation. Have a native speaker of each target market audition the read before you ship it.

How much does an Eleven v4 voiceover cost?

On fal, read on September 29, 2026, Eleven v4 costs $0.08 per 1,000 characters and Eleven v4 Turbo costs $0.04 per 1,000 characters. By our arithmetic a marked-up 30-second ad script of about 500 characters costs about 4 cents per take on v4 and about 2 cents on Turbo. On ElevenLabs' own plans, the company says v4 uses the same credit pricing as its other text-to-speech models.

Is Eleven v4 available in Novoads?

No. Novoads does not offer Eleven v4. In Novoads the voice is one layer of the video: a talking actor delivers your script with an AI voice, lip-sync and captions in the same workflow, with voices in 31 languages and their regional accents. Plans are listed on novoads.ai/pricing.

Key Takeaways

  • ElevenLabs launched Eleven v4 and the low-latency Eleven v4 Turbo on September 28, 2026, after previewing v4 in Warsaw in June. For ad voiceovers, v4 is the model to audition; Turbo is built for agents.
  • The #1 Artificial Analysis rank, the ~75% blind-test preference and both latency figures are ElevenLabs' own claims. Turbo's ~100 ms is median inference latency and its ~150 ms is median time to first speech: two different metrics.
  • Direction moved into the script. Plain-language delivery notes and inline tags like [laughs] or [pause] shape the read, and ElevenLabs says v4 follows them more accurately than prior models.
  • One voice across 90+ languages with a native accent is the localization claim to test first, with a native speaker per market. Clones made before v4 need retraining, and every clone needs verified consent from the voice's owner.
  • At fal's September 29 prices a 30-second take costs a few cents, so the voice is the cheapest line in the ad. The real cost is the listening, which is why a fixed test protocol beats auditioning by feel.
Mauricio Valdivia

Mauricio Valdivia

Founder of Novoads

Mauricio is the founder of Novoads, where he works to democratize video advertising with AI for brands in Latin America.

Related Articles

Gemini 3.8 Flash TTS for Ad Voiceovers: 30-Second Voice Replication, Consent First

Gemini 3.8 Flash TTS for Ad Voiceovers: 30-Second Voice Replication, Consent First

Google's September 23 launch adds prompt-designed voices, a library it puts at 2,000-plus voices, and replication of your own voice or one you have the rights to use from a short sample, gated by a spoken consent recording from the voice owner. Here is what that means for brand voices and spokesperson ads, and what a read costs.

NewsIndustry newsText-to-speech
The New Top-Ranked Text-to-Speech Model Ships Two Voices. What Ad Voiceovers Actually Need

The New Top-Ranked Text-to-Speech Model Ships Two Voices. What Ad Voiceovers Actually Need

Alibaba's new speech model took the number one spot on Artificial Analysis' Speech Arena for provider voices at an Elo of 1,238, yet its published catalog is two system voices in Mandarin and English with Spanish still missing from the shipping language table, which is exactly the distance between winning a blind listening test and being usable for a Spanish-language ad.

NewsIndustry newsText-to-speech
Open-Source ElevenLabs Alternatives for Ad Voiceovers: 6 Honest Picks and Their Real Tradeoffs

Open-Source ElevenLabs Alternatives for Ad Voiceovers: 6 Honest Picks and Their Real Tradeoffs

Open and local voice models now do a lot of what ElevenLabs does. Here are six real alternatives for ad voiceovers, compared on what actually decides an ad: commercial licensing, batch consistency, emotional range, and setup overhead.

ComparisonsElevenLabsOpen source
How to Keep One AI Actor Consistent Across Every Ad Variation

How to Keep One AI Actor Consistent Across Every Ad Variation

Actor drift is a workflow problem, not a prompting problem. Lock one reference still, drive it with a video you already approved, and vary one axis at a time. Here is the full production sequence, including the settings that quietly break a set halfway through.

GuidesAI actorsMotion control

Ready to create video ads with AI?

Generate professional video ads in minutes, not weeks.

Start for $49/month