Gemini 3.8 Flash TTS for Ad Voiceovers: 30-Second Voice Replication, Consent First
Google's September 23 launch adds prompt-designed voices, a library it puts at 2,000-plus voices, and replication of your own voice or one you have the rights to use from a short sample, gated by a spoken consent recording from the voice owner. Here is what that means for brand voices and spokesperson ads, and what a read costs.
Mauricio Valdivia
·11 min

Thirty seconds of audio, and one sentence of consent
A skincare founder wants every ad variant in her own voice. She has no studio. She has a phone, a script, and fifty hooks to test.
On 23 September 2026 Google introduced Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, calling them its "most expressive audio generation models yet." One line in the announcement reads as if it were written for her: the model can recreate a consistent vocal profile from a 30-second audio sample of your voice or a voice you have the rights to use.
The condition in that sentence is the story. Replication runs behind a consent check. Google says users must provide a verbal consent recording from the voice owner that matches the reference speaker before a voice can be created. Our reading, for anyone who makes ads: a brand voice just became cheap to produce and deliberately hard to take. This post reads the launch for ad voiceovers, covering what the two models do, what the consent gate means when the voice belongs to a creator or a spokesperson, what a read costs at Google's list price, and where the models run.
What Google shipped on 23 September
Two models, one set of controls, split by job. The launch at a glance:
- Models: Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS
- Voices: a library Google puts at 2,000+, plus voices designed from a text prompt
- Replication: your own voice or one you have the rights to use, from a 10 to 30 second clip plus a consent recording
- Direction: line-by-line stage directions and inline vocal tags
- Where: the Gemini API and Google AI Studio now, Gemini Enterprise later
Flash TTS for direction, Flash-Lite for volume
Google positions the pair by workload. Flash TTS is "built for deep creative direction and character design." Flash-Lite TTS is "built for high-volume, cost-efficient scale," optimized for high-volume dubbing, audio content creation and expressive voice agents.
The developer documentation sharpens the split into a "when to use which model" guide:
- Use Flash TTS when maximum acoustic fidelity, nuanced acting and expressive control are the top priority, including difficult pronunciations and regional or minority dialects.
- Use Flash-Lite TTS when the job is high-volume bulk production, read-aloud features, reliable voice replication or everyday single-speaker speech across major languages.
Both models share the exact same API schema and prompting format, so moving a job from one to the other is a single parameter.
For an ad pipeline that split maps cleanly onto how creative actually gets made. The hero read for the brand film goes to Flash TTS. The fifty hook variants you will kill by Friday go to Flash-Lite.
Where you can use it today
Per the announcement, the rollout started on 23 September and splits by audience:
- Developers: both models in the Gemini API and Google AI Studio
- Enterprises: API access in Gemini Enterprise, listed as "coming soon"
- Everyone else: Flash TTS in Gemini Notebook, Flash-Lite TTS in Google Vids
Google also says it is partnering with Figma, HeyGen, Linguana, Wondercraft, 99.co and Ollang, companies it describes as integrating the new models. Our read is that most ad makers will hear these voices first inside a tool they already pay for, not through the API.

Three ways to get a voice, and what each is for
The announcement frames the launch as a move from 30 original voices to "an infinite library." In practice there are three doors, and each one suits a different kind of brand:
- The library: fastest to a usable read, shared with every other buyer
- Voice design: a voice nobody else has, described in words
- Replication: your own voice or one you have the rights to use, with the owner's consent
Pick one from the library
The announcement says the library holds 2,000+ production-ready voices with broad language coverage, and names Mexican Spanish, Quebec French and Scots English as regional varieties in it. The developer documentation describes the same shelf in different words: 30 curated prebuilt studio voices, plus an Extended Voice Library of "hundreds of additional voices across languages, accents, and character archetypes."
We will not reconcile those two counts on Google's behalf. The practical advice is the same either way: if a campaign depends on one specific accent, list the voices through the API and listen before you plan around it. A number on a launch page is not an audition.
Describe one in a prompt
Voice design is the brand-voice feature for anyone without a voice to copy. Google says Flash TTS creates voices from scratch by customizing role, accent and voice characteristics across more than 100 languages and dialects with natural-language prompts. The docs list over 130 languages for Flash TTS and over 100 for Flash-Lite TTS.
A designed voice is meant to be reused. The announcement says saved custom voices keep consistent performance with minimal drift across ongoing projects, which is the audio version of the problem we cover in keeping one AI actor consistent across ad variations. The docs also set limits: stored voices, designed or replicated, are capped at 200 per project with a one-year time-to-live, and a stateless voice key lasts seven days.
Read that as a production rule. A brand voice on this API is an asset with an expiry date, so keep the prompt or the reference audio that produced it, the way you keep the master file of a logo.
Replicate one you have the rights to
The third door is the one in the headline. Google's announcement says Flash TTS can recreate a consistent vocal profile from a 30-second sample of your voice or a voice you have the rights to use, and the voice replication docs say both 3.8 TTS models support it. A fourth option, voice remixing, would let you fine-tune the timbre, pitch, pace and accent of a library voice. The announcement marks it "coming soon."
How voice replication works, consent included
Replication is where the launch stops being a spec sheet and becomes a workflow with a person in it. Everything below applies only to your own voice or a voice you have the rights to use.
Two recordings, not one
Google's voice replication documentation says every request needs two real human audio recordings from the same adult speaker:
- The reference: a 10 to 30 second clip of clean, natural speech from the speaker whose voice you want to replicate
- The consent clip: the same speaker reading Google's consent statement aloud
So the announcement's 30 seconds is the top of the accepted range, not a requirement. In English, the statement the speaker must read is:
I am the owner of this voice and I consent to Google using this voice to create a synthetic voice model.
The announcement adds the check that makes the pair meaningful: the verbal consent recording has to come from the voice owner and match the reference speaker before a voice can be created.
Thirty consent locales
The docs say the consent audio must recite the exact statement in one of 30 supported language locales. Spanish (Spain), Spanish (US) and Portuguese (Brazil) are among them. Google's recording advice is practical: minimize room echo, background noise, music and overlapping voices, and record both clips on the same microphone in the same acoustic setting so the speaker verification check succeeds reliably.
Turned into a session checklist, that advice becomes:
- Book one sitting and record both clips on the same mic, in the same room.
- Keep the reference between 10 and 30 seconds of clean speech, with no music under it.
- Have the speaker read the consent statement in a supported locale.
- File the two raw recordings with the signed agreement that covers the voice.
Where replication is switched off, and what gets stamped
A footnote on the announcement says voice replication through AI Studio is not available in Illinois, Texas, the EEA, the UK, Switzerland and India. That line is scoped to AI Studio, so check it against wherever your team actually works before you book the session.
The safeguards Google lists around replication, and one the docs add:
- Consent verification: built in, as described above
- SynthID: every audio clip generated by Google's Gemini Audio models is watermarked with it
- C2PA credentials: named in the announcement alongside the watermark
- Deletion: stored replicated voices can be listed, inspected and deleted through the Voices API
That last line matters for the next section.
What the consent gate means for a spokesperson ad
This is the part the headline number hides. Thirty seconds is easy. The sentence is the work.
Your own voice is the easy case
A founder replicating her own voice records two clips and is done. She owns the voice, she reads the statement, and the consent check passes because the speaker in both clips is the same person. For a solo brand this is the cleanest use of the launch: one recording session, then every future script read in the founder's voice without booking another. Google's own wording puts this case first, "your voice," before the rights-holder case.
A creator or a hired spokesperson
Here the condition does real work. You can only replicate a voice you have the rights to use, and the docs require both recordings to come from the same adult speaker. So the creator has to sit at the microphone and say the consent sentence in their own voice at least once. You cannot do it for them from a folder of their old reads.
Our stance: this moves the hard part of a replicated voice from the recording to the rights conversation, and that is where it belonged all along. Put voice replication in the creator agreement before the session, and make it name four things:
- Scope: what the replicated voice may say, and what it may never say
- Channels: which ad platforms and placements the voice can run on
- Term: how long the brand may keep generating with it
- Deletion: when the stored voice is deleted, and who confirms it
It is the same discipline that keeps an AI testimonial video honest: the words and the person have to be real, and the paperwork has to say so.
A voice you only have a file of
An old customer testimonial. A podcast guest who praised the product. A celebrity clip. As Google documents it, the consent recording has to match the reference speaker, so a file without its owner gives you a reference clip and nothing else.
That is the right design, and ad makers should want it. The same gate that stops you from lifting a voice stops a competitor from lifting your founder's.

Directing the read, line by line
A voice is only half a voiceover. The other half is the read, and this is where the launch is most useful for short ads.
Stage directions that stay silent
The docs say Gemini 3.8 TTS treats the text field strictly as a verbatim transcript, and split a performance into two places so the directions are not read aloud:
- The style field: sustained delivery across a whole turn, such as emotion, pacing and volume
- The transcript: the exact words, plus inline tags for momentary sounds
The announcement frames the same idea as directing performance line by line: write your own stage directions or let Gemini steer delivery with natural script cues.
Anyone who has watched a generator speak its own instructions out loud in a finished ad knows why that separation matters. The script is what the viewer hears. Everything else stays in the margin.
Tags for the small human noises
Momentary sounds go inline, in angle brackets inside the transcript. The docs' examples include <sigh>, <cough> and <short pause>, and the announcement also mentions laughs, gasps and active-listening interjections. A hook split the way the docs describe looks like this:
style: warm, a little surprised, quick
text: I did not expect a sunscreen to fix my makeup. <short pause> It did.
For a UGC-style read, one well-placed pause before the payoff does more than any adjective in the style field. Keep the style string short and constant across a batch, and let the script carry the variation you are actually testing.
Two voices in one script
The announcement lists native two-speaker scene staging. The docs add a limit: single-request multi-speaker generation supports up to two speakers using prebuilt voices, and designed or replicated voices in a dialogue must be synthesized one turn at a time. A two-person skit with your brand voice in it is therefore several calls stitched together, not one.
What comes out
A single, non-streaming request returns a complete WAV file at 24 kHz, mono, 16-bit PCM by default. That is audio only. The picture, the lip-sync and the captions are still separate jobs; we looked at the caption side in why a new speech-to-text model matters for ad captions.
What a voiceover costs at Google's list price
Every figure here comes from Google's Gemini API pricing page, last updated 24 September 2026, standard paid tier, audio output. None comes from the announcement.
| Model | Per 10 s, through 2026 | Per 10 s, from 2027 | 50 reads of 30 s, 2026 |
|---|---|---|---|
| Flash TTS | $0.00225 | $0.0045 | about $0.34 |
| Flash-Lite TTS | $0.0015 | $0.003 | about $0.23 |
The page prices audio at 25 tokens per second. Behind those per-second figures sit per-million-token rates of $9.00 for Flash TTS and $6.00 for Flash-Lite TTS through 31 December 2026, rising to $18.00 and $12.00 from 1 January 2027.
A worked example: fifty hook variants
The last column is our arithmetic, not Google's. A 30-second read is three 10-second units. On Flash TTS that is 3 × $0.00225, or $0.00675 of audio output per read. Fifty variants come to about 34 cents.
On Flash-Lite TTS the same read is 3 × $0.0015, or $0.0045, and fifty variants come to about 23 cents. We left out text input, which Google prices at $0.50 per million tokens through 2026; for a 75-word script that is a rounding error. From January 2027 both output rates double, so the same fifty reads cost about 68 and 45 cents.
The price is not the constraint
At these rates the voice is the cheapest line in the ad. What decides whether a replicated voice ever ships is everything around it:
- The rights to the voice, settled before the session
- The picture it sits on, which the voice model does not make
- The test budget that tells you which of the fifty hooks works
If the per-second price was your reason to look at self-hosted voices, the questions we raised about commercial licensing in open-source ElevenLabs alternatives still apply here, in a different shape.
Where it sits among the voice tools ad makers use
A launch this broad invites a leaderboard comparison. For an ad maker the better comparison is on terms, provenance and fit.
Launch specs versus usable voices
In July we read Alibaba's speech model the same way, in what ad voiceovers actually need from a top-ranked TTS model: the gap between a headline spec and the voices a buyer can actually use. The same test applies here, before a launch page becomes a plan:
- Count the voices in your buyer's language and accent, by listening, not by the headline number.
- Direct a real script, your actual hook with its pause and payoff, not a demo line.
- Check where replication is on for the surface your team will use, and who has to record the consent clip.
- Price the batch, not the clip: fifty variants, then the 2027 rate.
Provenance as the pitch
We read Adobe's Firefly audio launch through its licence terms in Adobe Firefly's AI audio tools for ads. Google's pitch here leans on consent verification, SynthID watermarking and C2PA credentials instead. A licence and a provenance trail are different promises, and a brand choosing a voice vendor should read both kinds of fine print.
Watermarks are not labels
A watermark makes AI audio detectable. It does not answer an ad platform's disclosure rules, which are their own question with their own answers per platform. We keep those in do you have to label AI-generated ads and in the three platform rulebooks compared.
How Novoads handles the voice in a UGC ad
Novoads does not run Gemini TTS, and nothing in this post is a feature you can select in the product. Novoads' talking actors speak from an ElevenLabs-backed voice catalog covering 31 languages with their regional accents.
The workflow is built around the finished ad rather than the audio file:
- Write or generate the script.
- Pick an AI actor and a voice.
- Get a vertical video ad with the lip-sync and captions already done.
If you want to hear your hook on a talking actor before you think about replicating your own voice or one you have the rights to use, you can try it in Novoads, and the plans are on the pricing page.

The voice got cheap. The permission did not.
Google priced a 30-second read at well under a cent and then put the voice owner's own sentence in front of every replicated voice, whether it is yours or one you have the rights to use. That ordering is the lesson for anyone making ads. The voice is a commodity now; the consent is the asset. Settle the agreement first, and the voice is the easy part.
Frequently Asked Questions
What are Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS?
They are two text-to-speech models Google announced on 23 September 2026. Google positions Flash TTS for deep creative direction and character design, and Flash-Lite TTS for high-volume, cost-efficient scale. Both are rolling out to developers in the Gemini API and Google AI Studio, with API access in Gemini Enterprise listed as coming soon. For everyone else, Flash TTS is rolling out in Gemini Notebook and Flash-Lite TTS in Google Vids.
Can Gemini TTS clone my voice for ads?
Google's announcement says Flash TTS can recreate a consistent vocal profile from a 30-second audio sample of your voice or a voice you have the rights to use. The API documentation asks for a reference clip of 10 to 30 seconds and a second recording in which the same adult speaker reads Google's consent statement, and it lists both 3.8 TTS models as supporting replication. A footnote on the announcement says replication through AI Studio is not available in Illinois, Texas, the EEA, the UK, Switzerland and India.
Can I replicate a creator's or spokesperson's voice?
Only a voice you have the rights to use, and only with that person's own consent recording. Google says the verbal consent recording must come from the voice owner and match the reference speaker before a voice can be created. In practice the creator has to record the consent statement themselves, so the agreement about what the voice may say, where and for how long should be settled before the recording session.
How much does Gemini TTS cost per voiceover?
On Google's Gemini API pricing page, last updated 24 September 2026, standard-tier audio output costs $0.00225 per 10 seconds on Flash TTS and $0.0015 per 10 seconds on Flash-Lite TTS through 31 December 2026. From 1 January 2027 those become $0.0045 and $0.003. By our arithmetic a 30-second read costs well under one cent of audio output on either model.
How many voices and languages does it support?
The announcement says the library holds 2,000+ production-ready voices and that voice design works across more than 100 languages and dialects. The API documentation describes 30 curated prebuilt voices plus an Extended Voice Library of hundreds of additional voices, and lists over 130 languages for Flash TTS and over 100 for Flash-Lite TTS. If a campaign depends on one specific accent, audition it before planning around it.
Is Gemini TTS available in Novoads?
No. Novoads does not run Gemini TTS. Its talking actors speak from an ElevenLabs-backed voice catalog covering 31 languages with their regional accents, and the voice, lip-sync and captions are produced in the same video workflow.
Key Takeaways
- Google's 23 September 2026 announcement says Flash TTS can recreate a consistent vocal profile from a 30-second sample of your voice or a voice you have the rights to use. The API docs accept a 10 to 30 second reference clip, so 30 seconds is the ceiling, not the requirement.
- Replication needs a second recording: the same adult speaker reading Google's consent statement in one of 30 supported locales. The voice owner becomes a participant in the session, not a file on your drive.
- Direction lives outside the script. The docs say the text is read as a verbatim transcript, sustained delivery goes in a style field, and inline tags handle pauses and sighs.
- At Google's dated list price a 30-second read costs well under one cent of audio output through 2026, and the rate doubles on 1 January 2027. The voice is the cheapest line in the ad; the rights are the expensive one.
- Novoads does not run Gemini TTS. Its talking actors use an ElevenLabs-backed catalog in 31 languages, so this launch is a reason to watch the voice market, not a feature you can select in the product.




