Skip to main content

The New Top-Ranked Text-to-Speech Model Ships Two Voices. What Ad Voiceovers Actually Need

Alibaba's new speech model took the number one spot on Artificial Analysis' Speech Arena for provider voices at an Elo of 1,238, yet its published catalog is two system voices in Mandarin and English with Spanish still missing from the shipping language table, which is exactly the distance between winning a blind listening test and being usable for a Spanish-language ad.

Mauricio Valdivia

Mauricio Valdivia

·10 min

The New Top-Ranked Text-to-Speech Model Ships Two Voices. What Ad Voiceovers Actually Need

Rank one on the board, two voices in the catalog

A Spanish-language advertiser reads the headline, opens a second tab, and goes looking for the pricing page. That instinct is correct. It will also cost an afternoon.

On 20 July 2026 Alibaba's Tongyi Lab released Qwen-Audio-3.0-TTS, and two days later its Plus tier sat at the top of Artificial Analysis' Speech Arena leaderboard for provider voices with an Elo of 1,238, ahead of Speechify's Simba 3.2 at 1,229 and Google's Gemini 3.1 Flash TTS at 1,211. That is a real result on an independent board nobody at Alibaba controls, measured by blind listener votes rather than by a vendor's own eval harness. It is also, for anyone shipping ads in Spanish or Portuguese, close to unusable on its own. The catalog published behind that number is two system voices, and both are listed as Mandarin and English.

What Alibaba actually shipped on 20 July

The launch is one model line, two API tiers, and a control layer that is more interesting than the ranking that made the news. The shape of it, as Alibaba documents it:

  • Tiers: Flash for real-time interaction, Plus for high-quality generation
  • Languages: 16, with 20 Chinese dialect regions listed as a separate spec
  • Control: free-style natural-language instructions plus 86 inline tags
  • Long form: one-pass synthesis up to 3 minutes
  • Access: Alibaba Cloud Model Studio only, over WebSocket or the DashScope SDK
  • Weights: none published, this is an API product

Two tiers from one lineage

Tongyi Lab describes the release as two variants from the same lineage. Flash is "tuned for real-time interaction, with a first-packet latency at 300ms-level", aimed at voice agents and anything conversational. Plus is tuned for high-quality generation, where naturalness and timbre fidelity matter more than speed. The tier that tops the leaderboard is Plus, so every number in the arena describes the slower, quality-first variant.

A hosted API, not a download

This is not an open-weights release. The model ships through Alibaba Cloud Model Studio as two SKU ids, qwen-audio-3.0-tts-plus and qwen-audio-3.0-tts-flash, called over a WebSocket connection or the DashScope SDK, with the documentation listing China (Beijing) and Singapore as the regions. If your stack does not already terminate in Alibaba Cloud, adopting this model is an infrastructure decision before it is a creative one. That is a different proposition from the wave of open-weight voice models you can host yourself.

It is also the second closed launch from the same lab in the same week. The image side went further in the same direction: Qwen Image 3 arrived with no weights, no benchmark table and no technical report, breaking with a series that had previously shipped Apache 2.0 weights and a same-day paper. Read together, the two releases say something about where this lab is heading that neither says alone.

The control layer is the real headline

The model page describes free-style natural-language instruction following alongside 86 newly added inline tags, and Alibaba Cloud's docs confirm the shipping surface: the service "Supports Instruction control, which lets you control speech expressiveness through natural language instructions", exposed as an instruction parameter. Both tiers carry it. That, not the Elo, is the part worth a serious look.

A real UGC creator filming a product testimonial on a phone
Novoads · UGC video ads with AI, ready in minutes.
Try now

How to read an Elo of 1,238

Arena scores are the most quotable and most misread numbers in AI. Three details change what this one means.

The gap that is not a gap

Artificial Analysis publishes a 95% confidence interval next to every score. The leader's row reads 1,238 with an interval of plus or minus 16, on 1,479 samples. Simba 3.2 reads 1,229, also plus or minus 16. Those intervals overlap, which means the honest reading of the top of that board is a tie, not a victory. Artificial Analysis says as much in its own layout by giving both models a rank range of one to two. The number is also live: the same board read 1,236 the day before, because votes keep arriving. Any post that quotes an arena Elo without a date is quoting a snapshot and calling it a fact.

Which arena this is

The board in question compares models using each provider's own native voices. Artificial Analysis runs a second one that compares models on "the same 8 cloned voices (4 US, 4 UK)", which is much closer to what a brand does when it commits to a single spokesperson voice across a campaign. That board has a different leader, Cartesia's Sonic 3.5 at 1,108, with ElevenLabs' Eleven v3 second at 1,079. Two boards, two answers, and the headline only travelled with one of them.

What a blind vote does not test

An arena vote asks one question: which of these two short samples sounds better. The questions that actually decide an ad voiceover are elsewhere.

  • Does a voice exist in your buyer's language and regional accent?
  • Does it hold together across a 45-second script, not a 10-second sample?
  • May you ship the output in a paid ad, from the region you operate in?
  • Can you direct it to sound like a person recommending a moisturizer rather than a narrator reading a paragraph?

None of those are on the scoreboard, and the first one is where this particular launch gets complicated.

The language question a Spanish-speaking advertiser has to ask

This is where the story stops being about a ranking. If your market speaks Spanish or Brazilian Portuguese, the only question that matters is whether the model will read your script in a voice your buyer recognizes as local.

What the 16-language list contains

Tongyi Lab enumerates them by name: "Arabic, Chinese, English, French, German, Indonesian, Italian, Japanese, Korean, Malay, Portuguese, Russian, Spanish, Tagalog, Thai, Vietnamese." Spanish and Portuguese are both there. The list names languages, not locales, so nothing on that page distinguishes Latin American Spanish from peninsular Spanish, or Brazilian Portuguese from European Portuguese. Alibaba's own numbers, self-reported in the same post, claim the best word and character error rates in 10 of the 16.

What the shipping docs list today

Alibaba Cloud's own product documentation tells a narrower story. Three findings from the shipping pages, checked on 22 July 2026:

  • System-voice languages for the Qwen-Audio-TTS series: Chinese (Mandarin) and English
  • Cloned-voice languages: Chinese plus English, Japanese, Korean, German, French, Italian, Russian, Portuguese, Thai, Indonesian, Malay and Vietnamese. Portuguese is on that list. Spanish is not
  • The published voice list for the Plus tier: two voices, each with a Language field reading Chinese (Mandarin), English

Coverage is also a per-voice property rather than a model-wide one, and the docs say so directly, instructing callers to "select a voice that supports the target language". A model-level count of 16 therefore tells you almost nothing about whether the voice you pick will read your script.

The asterisk that reconciles them

These two pages are not in contradiction, because Alibaba flagged the gap itself. The 16-language list carries an asterisk: "*Full language support rolling out soon." That single line is the most useful sentence in the whole launch for a Spanish-language advertiser. The capability is announced, the surface has not caught up, and the honest status as of 22 July 2026 is that you cannot select a Spanish voice on the tier that won the board. Everything else is a forecast.

Directing the read is worth more than winning the blind test

Strip out the ranking and there is still a genuine advance here, and it is the one closest to how ad scripts actually get made.

Instructions instead of parameters

The model page says it "interprets free-style natural-language instructions describing role, emotion, speaking style, rate, timbre, and accent". Tongyi Lab frames the same feature in practitioner terms: describe the delivery you want in plain language instead of hand-tuning acoustic parameters. Their own example prompt is a stadium announcer with broad pacing and a lifted intonation on the welcome. That is a direction you would give a human voice actor, typed into an API field.

Tags for the beats between words

Alongside the free-form instruction, the release adds "86 newly added fine-grained inline tags" that sit inline in the script, marking a gasp, a laugh, a shift in tone at a specific word rather than across a whole take. Ad reads live and die on those beats. The half-laugh before "honestly, I didn't expect it to work" is the thing that makes a testimonial sound like a person.

Why this maps onto ad scripts

Every UGC-style script is already a delivery instruction wearing a costume. The reason AI testimonial videos fall flat is almost never the timbre of the voice; it is that the read is uniformly warm, uniformly paced, and uniformly wrong for a hook that needs to sound slightly annoyed. Per-phrase control is the lever that closes that gap, and it is the reason this launch matters more than its Elo does.

UGC creators each holding a different product up to the camera
Novoads · UGC video ads with AI, ready in minutes.
Try now

Three tradeoffs a voiceover pipeline feels first

Quality is the fun number. These three are the ones that show up in a Tuesday-afternoon batch of 20 variants.

Throughput

Artificial Analysis measures the model at 16 characters per second. For comparison, in the same measurement Sonic 3.5 runs at 120 and Simba 3.2 at 30.2. Voiceover work is not one clip, it is a rack of script variants regenerated every time a hook changes, so generation speed compounds in a way it never does for a single demo.

Put real numbers on it. Novoads sizes a script at 15 characters per second of speech, so a 45-second read is 675 characters. A normal test cycle is 20 variants of that read, which is 13,500 characters. Divided by the measured 16 characters per second, one batch is about 844 seconds of pure generation, and that clock runs again on every rewrite. On Simba 3.2's measured 30.2 it is roughly half that; on Sonic 3.5's 120 it is a fraction. Quality-first tiers are usually slow tiers, and this one is unusually so.

Price per million characters

Artificial Analysis lists API prices in a single comparable unit, the cost to generate 1M characters on the creator's API at default settings. On that board the new leader is cheaper than most of what it beats:

ModelArena EloListed price per 1M characters
Qwen-Audio-3.0-TTS-Plus1,238$27.6
Simba 3.21,229$10.0
Gemini 3.1 Flash TTS1,211$18.3
Sonic 3.51,209$49.0
Eleven v31,172$100.0

A million characters is a lot of ad copy, so for most advertisers none of these rows is the deciding cost in a campaign. The row that should catch your eye is the second one: the model in a statistical tie for first place is listed at well under half the price.

Which tier can clone a voice

Here is the detail that reframes the whole launch. On Alibaba Cloud's own model table, qwen-audio-3.0-tts-plus is marked as supporting instruction control while voice cloning and voice design are both unsupported. Cloning lives on the Flash tier. So the variant that topped the provider-voice board is the preset-voice variant, and a brand that wants its own spokesperson voice would be using the other one, which is not the model being ranked. Alibaba is explicit about the competitive framing on that same page, telling teams "currently using ElevenLabs, OpenAI, or Google for speech synthesis" which Model Studio model to migrate to. It is a positioning statement from a vendor, and it is worth reading as exactly that.

How Novoads handles ad voiceovers today

To be direct about what this news does not change: Qwen Audio is not in Novoads, and one leaderboard result is not a reason to put it there.

The voice layer we actually run

Novoads generates the voice for a talking-actor ad from an ElevenLabs-backed catalog spanning 31 languages and their regional accents, which is the layer that decides whether a Chilean or Mexican or Brazilian viewer hears someone from their own market or a neutral dub. That coverage is the single reason we have not chased a benchmark leader. A voice engine that scores well on English and Mandarin samples solves a problem we do not have, and does not yet solve the one we do.

What voice costs inside a video

The pricing is boring on purpose. AI voice is billed at 0.9 credits per minute of generated audio, and Voice Design, which turns a written description of a voice into previews, costs 0.5 credits per request. Those sit inside the same credit balance as the video, so a script rewrite that changes the read does not open a second vendor bill. When you generate an ad in Novoads, the script, the actor, the voice and the captions come out of one flow.

Why a benchmark win is not a migration

We have run this evaluation before, in public, on the transcription layer beneath auto-captions and on AI music for ad soundtracks. The pattern repeats: a model posts a headline number, the number is real, and the production question turns out to be about coverage, control, and commercial terms instead.

The right response is a bake-off on your own material, not a migration announcement. Four things worth testing before any voice engine gets near a live campaign:

  • Your actual script, in your actual market's Spanish or Portuguese, not the vendor's demo sentence
  • A repeat generation of the same line, because consistency across variants is what makes a batch usable
  • The awkward words: your brand name, a price, a product name nobody has ever transcribed
  • A directed read, where you ask for annoyed, or rushed, or amused, and check whether the delivery actually moves
Novoads UGC ad templates gallery
Novoads · UGC video ads with AI, ready in minutes.
Try now

A leaderboard measures a demo. A campaign measures a market.

There is a version of this story that reads "new best voice model, everyone else is behind," and it would be defensible on the evidence and useless in practice. The version that survives contact with a real ad account is narrower and more interesting: an independent board now ranks a Chinese-built, instruction-directed model first among provider voices, at a listed price roughly a quarter of ElevenLabs' flagship, while its own vendor documents two English-and-Mandarin voices and marks broader language support as still rolling out. Both things are true on 22 July 2026, and only one of them affects what you can ship this week.

The lesson generalizes past this model. The voice layer under AI-generated UGC ads keeps improving in public, and every few months a new engine takes a benchmark crown. What decides your creative is not who holds the crown, it is whether the voice speaks your buyer's Spanish, whether you can direct the read line by line, and whether it arrives already wired to the avatar and the captions. Rankings move weekly. Distribution does not. If you want the voiced version of your next ad without assembling any of this yourself, Novoads generates the actor, the voice and the captions in one pass. Three days of full access for $1, then $49/mo if you keep it. Cancel anytime.

Frequently Asked Questions

What is Qwen-Audio-3.0-TTS?

It is a hosted text-to-speech model released by Alibaba's Tongyi Lab on 20 July 2026, sold through Alibaba Cloud Model Studio in two tiers. Flash is tuned for real-time interaction at a first-packet latency Tongyi Lab describes as 300ms-level, and Plus is tuned for high-quality generation. The model page states support for 16 languages and, as a separate spec, 20 Chinese dialect regions, along with one-pass long-form synthesis up to 3 minutes. It is an API product, not a downloadable set of weights.

Is it the best text-to-speech model right now?

It is ranked first on one board. As of 22 July 2026, Artificial Analysis' Speech Arena leaderboard for provider voices places it at an Elo of 1,238, ahead of Simba 3.2 at 1,229. Both rows carry a 95% confidence interval of plus or minus 16 Elo, so the two overlap and the top of the board is better described as a statistical tie than a clear win. Artificial Analysis also runs a separate board on the same eight cloned voices, and that one is led by Cartesia's Sonic 3.5.

Does it support Spanish and Portuguese?

Spanish and Portuguese are both named in Tongyi Lab's list of the 16 supported languages, and the same list carries an asterisk reading that full language support is still rolling out. On the shipping side the picture is narrower: Alibaba Cloud's model-selection page lists the series' system voices as Chinese (Mandarin) and English, its cloned-voice language list names Portuguese but not Spanish, and the published voice list for the Plus tier contains two voices, both Mandarin and English. Treat Spanish as announced rather than shipped until that documentation catches up.

Can you use it for a branded or cloned ad voice?

Not on the tier that tops the leaderboard. Alibaba Cloud's own model table marks voice cloning and voice design as unsupported for qwen-audio-3.0-tts-plus and supported for the Flash tier. So the version that won the blind listening test is the preset-voice version, and a branded voice would run on the other tier, which is not the one being ranked.

Is Qwen Audio available in Novoads?

No. Novoads generates ad voiceovers from an ElevenLabs-backed voice catalog covering 31 languages with their regional accents, exposed in the app as AI voice and Voice Design. Qwen Audio is not part of that stack and is not selectable anywhere in the product.

What should actually decide your voice engine for ads?

Four things a leaderboard does not measure: whether a voice exists in your buyer's language and regional accent, whether you can direct the read rather than just generate it, whether the commercial terms and the region you can call it from work for your business, and whether it plugs into the rest of your creative pipeline. An arena Elo is a preference score on short samples in a handful of voices, which is a real signal about quality and a weak signal about fit.

Key Takeaways

  • As of 22 July 2026 the model leads Artificial Analysis' Speech Arena for provider voices at an Elo of 1,238, but the nine-point gap to second place sits inside overlapping confidence intervals, so the top of that board is a statistical tie.
  • The 16-language claim and the shipping documentation disagree today. Tongyi Lab names Spanish and Portuguese among the 16 and marks the list as still rolling out, while Alibaba Cloud's published voice list for the winning tier is two voices, both Mandarin and English.
  • The genuinely new capability is direction, not timbre: free-style natural-language instructions covering role, emotion, style, rate, timbre and accent, plus 86 inline tags for the non-verbal beats.
  • The tradeoffs a batch pipeline feels first are throughput at 16 characters per second, an API reachable only through Alibaba Cloud Model Studio, and a Plus tier that cannot clone a voice at all.
  • Novoads runs an ElevenLabs-backed voice catalog across 31 languages and their regional accents. A model topping a blind-preference board is a reason to keep testing, not a reason to migrate a production ad pipeline.
Mauricio Valdivia

Mauricio Valdivia

Founder of Novoads

Mauricio is the founder of Novoads, where he works to democratize video advertising with AI for brands in Latin America.

Ready to create video ads with AI?

Generate professional video ads in minutes, not weeks.

Start for $1