What Is Lyria 3.5? Google's New Music Model and the Ad Soundtrack Gap
Google rolled out Lyria 3.5 in Flow Music on July 29, 2026, with advances in musicality, lyrics, vocals and easier tempo and duration control. Here is what a better music model does, and does not, change for the one layer of an AI ad that still gets outsourced to a stock library.
Mauricio Valdivia
·11 min

Google shipped a better music model, not a better pipeline
The ad is done. Everything except the music. Twenty-two seconds of a person holding a serum bottle, the voiceover timed to the product reveal, the captions burned in, and a silent gap where the bed is supposed to go. So somebody opens a stock library, sorts by "upbeat," auditions nine tracks that all sound like the same laptop commercial, picks the fourth one because it is 4pm, and ships.
That last step is the only part of the process a machine did not touch.
On July 29, 2026, Google announced Lyria 3.5, calling it its newest music generation model, and rolled it out the same day in Google Flow Music. Google's post lists four advances: richer melodic structures, higher quality lyrics with better prompt adherence, more expressive vocals with improved pronunciation, and easier control over tempo and duration. The DeepMind model card published the same day describes the model plainly as a music generation system that synthesizes high-quality audio from a text prompt.
Two things are worth saying before anything else, because most of the coverage blurs them. First, this is a music model, not a voice model. It sings; it does not narrate your ad. Second, the model card names one distribution channel, Google Flow Music, which makes Lyria 3.5 a place you go rather than a step your pipeline calls. Novoads does not run it.
So this is not a "new tool for your workflow" post. It is a read on what a genuinely better music model means for the last layer of an AI ad that nobody automated, and what to do about that layer this week.
What Google actually announced on July 29
The four advances, in Google's own words
Google frames Lyria 3.5 as its newest music generation model delivering significant advancements across musicality, lyrics and vocal quality. Underneath that headline sit four specific lines:
- Musicality. Richer, more complex melodic structures that sound more natural. The advertiser's read: fewer beds that audibly loop at the fifteen-second mark.
- Lyrics. Higher quality lyrics with improved prompt adherence and structural awareness. The read: a sung hook you can steer toward the product instead of away from it.
- Vocals. More realistic and emotionally nuanced vocals, plus improved pronunciation. The read: a vocal in your market's language that does not sound machine-read.
- Creative control. More easily control the tempo and duration of your outputs. The read: careful, because this is the one line Google states without publishing what the control actually is.
That fourth line is the one everybody quoted, and it deserves a caveat. It appears in the launch blog post and nowhere in the technical documentation. Google says control got easier; Google does not publish what the control surface is. Press coverage has filled that vacuum with specifics that no Google page states, and those specifics are worth ignoring until a Google surface carries them.
What the model card adds, and where it is quieter
The model card is a different document written for a different reader, and it is more useful than the blog post if you care about what shipped rather than how it feels.
It confirms the shape of the thing: inputs are text, outputs are audio in the form of music plus text in the form of lyrics. On Google's own evaluation, Lyria 3.5 improved significantly compared to Lyria 2 on audio fidelity, and with lyrics it demonstrates better prompt adherence, following both simple and more complex instructions more accurately.
Sort the launch by which Google document carries each claim, and a pattern shows up:
| The claim | Where Google states it | Kind of statement |
|---|---|---|
| Richer melodic structures | Launch post | Marketing |
| Better lyric prompt adherence | Launch post and model card | Evaluation |
| Improved pronunciation | Launch post | Marketing |
| Audio fidelity gain over Lyria 2 | Model card | Evaluation |
| Easier tempo and duration control | Launch post | Marketing |
| SynthID watermarking | Model card and Lyria page | Mitigation |
Musicality and vocals get marketing sentences. Lyrics and fidelity get evaluation sentences. And the row a working advertiser should care most about, the watermark, never appears in the announcement everyone read.
On training data, the card says Lyria 3.5 was trained on audio data, and that audio datasets were annotated with text captions at different levels of detail. That is the whole disclosure. If you have been following the litigation around AI music training sets, you already know why a sentence at that altitude is not a chain of title, and why an advertiser cannot treat it as one.
The one channel it ships through
The distribution section of the model card names Google Flow Music. The Lyria family reaches further than that, and DeepMind's Lyria page points at several places to try it, but the model card's channel line for 3.5 specifically is the one that matters for anyone planning a workflow.

Why the soundtrack is the last layer handed back to you
Everything above the audio track got automated first
Look at what an AI ad pipeline now does with nobody touching it:
- Script. Drafted from a product page or a one-line brief, in the language you chose.
- Actor. A person who does not exist, delivering that script to camera and holding the product.
- Product imagery. Generated from one uploaded photo, at whatever angle the shot needs.
- Voice. Rendered straight from the script text, priced by characters rather than by takes.
- Assembly. Cuts, captions, and a reframe for six different platform spec sheets.
Then it stops, hands you a nearly silent file, and you go find music.
The music bed is where the assembly line stops
This is not an accident of tooling. It is a structural consequence of how music rights work. Video models train on video, and the resulting clip is a new frame sequence nobody recognizes. A music model produces something that sounds, by design, like music you have heard, in a market where the underlying catalogs are actively litigated and where a recognizable melody is the entire product. The legal surface is simply hotter, so vendors have been slower to wire music into pipelines than to wire in video, voice or images.
You can see the asymmetry in what shipped this year:
- Voice has real competitive depth, from open-source models you can host yourself to a text-to-speech model that topped the arena and still shipped two voices.
- Video models now generate synchronized audio in the same call as the picture, so the sound design arrives attached to the frames.
- Music sits outside all of that, in its own application, behind its own front door.
Batch math breaks at the bed
Here is the part that costs real money, and it only shows up at volume.
Take a normal week of creative testing: twelve variants of a twenty-two-second cut. At the fifteen-characters-per-second pacing Novoads uses to estimate read time, a twenty-two-second script is roughly 330 characters. Voice is billed at 1 centi-credit per 100 characters, so each variant's voiceover costs 4 centi-credits, which is 0.4 credits. Twelve variants of voice come to 4.8 credits, generated in one batch, with no human in the loop.
Now price the soundtrack for the same twelve. There are two options and both are bad:
- Reuse one bed across all twelve. You have introduced a variable shared by every cell in a test whose entire job is to isolate one variable. The bed is now a confound you built on purpose.
- Source twelve different beds by hand. Twelve trips to a library, twelve rounds of auditioning, twelve downloads, twelve filenames to keep straight. Call it an afternoon.
The generative layers of that ad cost single-digit credits and zero attention. The one non-generative layer costs an afternoon. That ratio is the actual story of a music model release, and it does not change until the model is callable from the thing that renders the ad.
What a better music model would fix, if you could reach it
Length that matches the cut, not the other way around
The oldest indignity of ad audio is trimming a full-length song down to twenty-two seconds and hoping the fade lands somewhere musical. DeepMind markets Lyria as generating cohesive tracks that match a project's exact duration, which inverts that. You would ask for the length of your cut and get a piece that resolves there, instead of an excerpt that just stops.
For a format built on tight cuts, that is a bigger quality-of-life change than better melodies.
Vocals in the language the ad is in
DeepMind also markets generating vocals in different languages across a wide genre range. Pair that with the launch post's pronunciation line and you get the thing that actually matters for anyone running ads outside English: a bed with a sung hook that does not sound imported.
Localized ad audio is currently a voiceover problem with a music-shaped hole in it. A track whose vocal is in the same language as the read is a different creative object than a track with an English hook under a Spanish script, and until now that has been a casting decision, not a prompt.
A watermark you can point at
This one did not appear in the launch post at all, and it is the most useful detail of the three.
The Lyria 3.5 model card lists SynthID watermarking among its product-level mitigations, and DeepMind's Lyria page states that all of its tracks are imperceptibly watermarked with SynthID technology, so that music created or edited with AI can be detected. Read that as an advertiser rather than as a musician. The uncomfortable question about a generated track has never been "will anyone notice," it has been "what will you say when someone asks where this came from." A documented watermark is a partial answer to the second question, and a partial answer beats the shrug that generated audio usually comes with.
It cuts both ways, obviously. Detectable means detectable by you and by everyone else. But a track you can describe beats a track you cannot, every time.

Access is the constraint now, not quality
A distribution channel is a workflow decision
Model quality and model reach are separate products, and the industry keeps confusing them. Lyria 3.5's quality claims are documented on two Google surfaces. Its reach, for this version, is documented as one channel: Google Flow Music.
For a person making one ad, that distinction is invisible. For a team making forty a month it is the whole thing:
- A model inside an app is a destination somebody visits. Output arrives by hand, one file and one browser tab at a time.
- A model behind an interface is a step something schedules. Output arrives with the rest of the batch, at three in the morning, unattended.
Everything that makes AI advertising economically interesting, which is variant volume, lives in the second category. Nothing in the first category scales past the patience of the person doing the clicking.
Manual export is a real workflow with a real ceiling
None of that makes the manual path worthless. Generate in Flow Music, export, drop it in the edit. For a hero video, a brand film, a launch spot you will run for six months, an afternoon of auditioning generated beds is a perfectly good use of an afternoon, and the output will likely beat the stock library.
The ceiling shows up at variant three. If the soundtrack requires a human to open a second application, the soundtrack is now the slowest step in a process explicitly built to remove slow steps. Teams that run structured creative operations will recognize this shape immediately: it is a manual dependency inside an automated cadence, and those either get eliminated or they set the pace for everything around them.
The contrast worth studying sits one layer over in video. On the render side, audio already arrives inside the same call that produces the picture: the animated spot walked through in 3D animated product ads without a studio comes back with its sound effects, ambient bed and lip-synced lines attached, because the model generates them in the same pass. That is exactly the property music is missing, and it is why the gap reads as a distribution problem rather than a quality one.
Google also publishes the honest caveat, which is rarer than it should be. DeepMind's own limitations note says the team is still working on improving key capabilities and that you should always carefully check that the tracks you create fit your vision. That is a reasonable thing to say about any generative audio in 2026.
What to do about your soundtrack this week
Default to a licensed bed, and keep the paperwork
The unglamorous answer is still correct, and there are only two versions of it:
- A production-music subscription carrying a commercial-use license, with the license file stored beside the creative rather than in somebody's inbox.
- Your ad platform's own cleared sound library, which is scoped to advertising use by design and needs no separate paperwork at all.
Both cost less than one hour of the conversation you would otherwise be having. The only habit worth adding is recording which one you used and when, because terms move and you want the version you agreed to.
Cut the bed entirely, which is usually the better ad anyway
Here is the unfashionable position, and I will not hedge it: most creator-style ads sound better with no music at all.
A bed is a polish signal. Polish is exactly what the UGC-style format is engineered to avoid, because a paid video that sounds produced stops reading as a recommendation and starts reading as an ad. Room tone, a phone microphone, a person talking slightly too fast. That is the texture the feed rewards, and a swelling indie-pop bed underneath it is the single fastest way to announce that a brand made this.
So the compliance-safe option and the performance option point the same direction. The soundtrack problem that a better music model would solve is, for a large share of ads, a problem you can decline to have.
When a bed genuinely earns its place
Not every ad is a talking head, so this is a judgment call rather than a rule.
Keep the bed when:
- No voiceover is carrying the timing, so the audio has to do the pacing.
- The piece is a montage, a product-motion cut, or a fashion film with no dialogue.
- The ad will run for months, so an afternoon of sourcing amortizes properly.
Cut the bed when:
- A person is talking to camera for most of the runtime.
- You are testing more than three variants of the same script.
- The ad is meant to land as a recommendation rather than as a commercial.
Where the bed does earn its place, a generated track with a documented watermark is a genuine upgrade over a stock cue you found in four minutes, and the manual export trip is worth taking.
How Novoads handles the layers that are actually generative
Novoads generates the parts of an ad that carry the message, from a script you write or auto-generate:
- The video, from an uploaded product image or a prompt.
- The AI actor, delivering that script to camera and holding the product.
- The voiceover, in voices spanning 31 languages.
It does not generate music, and that is a choice rather than a roadmap item. The music bed is the layer with the murkiest provenance and the smallest contribution to a creator-style ad, so we automate the layers that decide whether the ad works.
The credit math follows the same logic. Voice is billed by script length at 1 centi-credit per 100 characters, so a minute of read is 0.9 credits, and a one-minute talking-actor video is a flat 10 credits. Cheap enough that variant count stops being a budget question, which is the entire point of generating the layers you can. You can start for $1: $1 for 3 days of access, cancel any time.
If you do want a bed, add it in the edit from a licensed library or your platform's own cleared tools, and keep the receipt with the asset.

The model was never the ad's bottleneck
Lyria 3.5 is a real advance and Google documented it more carefully than most launches get documented, with a dated model card, a stated evaluation, and a published watermarking commitment. If you make music, that is good news on its own terms.
If you make ads, the news is narrower and more honest than the headline suggests. The reason ad soundtracks have been mediocre for a decade was never that the available music was not good enough. It was that clearing it was annoying, sourcing it was manual, and nobody could describe where it came from six months later. A better model fixes the first problem, which was not the problem.
Watch the channel line, not the capability list. The day a music model of this quality becomes a step a pipeline can call, with provenance attached, the soundtrack stops being the layer where the assembly line hands you a file and wishes you luck. Until then, the best soundtrack decision an ad team can make is still the boring one, and quite often it is silence.
Frequently Asked Questions
What is Lyria 3.5?
Lyria 3.5 is Google's newest music generation model, announced on July 29, 2026 and rolled out the same day in Google Flow Music. Google's DeepMind model card, published that day, describes it as a music generation system that synthesizes high-quality audio from a text prompt, taking text as input and returning audio plus lyrics. Google's announcement lists four advances: richer melodic structures, higher quality lyrics with better prompt adherence, more expressive vocals with improved pronunciation, and easier control over the tempo and duration of outputs.
Can Lyria 3.5 generate the voiceover for my ad?
No. This is the single most common mix-up around the launch. Lyria 3.5 is a music model, and Google's own model card is explicit that its outputs are audio in the form of music, plus text in the form of lyrics. It sings, it does not narrate. The voiceover layer of an ad is a separate class of model entirely, and it is where the language coverage and pronunciation control that ads actually need currently live.
Is there an API for Lyria 3.5?
The distribution section of Google's model card names one channel for Lyria 3.5: Google Flow Music. So the realistic workflow today is manual. You generate in Flow Music, export the file, and drop it into your edit by hand. That is a real workflow and plenty of people will use it well, but it does not batch the way the rest of an AI ad pipeline batches, and Novoads does not run the model.
Does Novoads generate music for ads?
No. Novoads generates the layers a UGC-style ad genuinely needs, which are the video, the AI actor and the voiceover. There is no music model in the catalog, and that is deliberate rather than a gap waiting to be filled. The music bed is the layer with the least clear provenance and the one a creator-style ad least needs, so we generate the parts that carry the message and leave the bed to a licensed source or to silence.
What is SynthID and why does it matter for an ad soundtrack?
SynthID is Google's watermarking technology for AI-generated media. The Lyria 3.5 model card lists SynthID watermarking among its product-level mitigations, and DeepMind's Lyria page states that all of its tracks are imperceptibly watermarked with SynthID so that AI-created or AI-edited music can be detected. For an advertiser that cuts both ways. It means a track you generate is identifiable as generated, and it means you have something to point at when someone asks where the audio came from, which is a question that normally gets a shrug.
What should I use for an ad soundtrack right now?
Three practical paths, in rough order of how often they fit. Use a licensed production-music library and save the license file next to the creative. Use your ad platform's own cleared commercial sound library, which is scoped to advertising use by design. Or build a voice-forward cut with no music bed at all, which is how most high-performing creator-style ads already sound. The third option removes the problem instead of managing it, which is why it is worth trying first.
Key Takeaways
- On July 29, 2026, Google announced Lyria 3.5, its newest music generation model, and rolled it out the same day in Google Flow Music with advances across musicality, lyrics and vocal quality, plus easier control over the tempo and duration of outputs.
- It is a music model, not a voice model. The DeepMind model card published that day describes it as a music generation system that takes text in and returns audio plus lyrics, so it does not touch the voiceover layer of an ad.
- The model card names Google Flow Music as the distribution channel. That makes Lyria 3.5 a place you go, not a step a pipeline calls, and Novoads does not run it.
- The genuinely useful detail for advertisers is not in the launch post. The model card lists SynthID watermarking among the product-level mitigations, and DeepMind's Lyria page says all of its tracks are imperceptibly watermarked, which gives an advertiser something to point at when someone asks where a track came from.
- Music quality was never the reason ad soundtracks are bad. Access and provenance were. Until a music model is callable from the tool that renders the ad, the practical moves stay the same: a licensed bed with the receipt saved, or a voice-forward cut with no bed at all.




