Skip to main content

Seedance 2.5 Prompt Guide for UGC Ads: Direct a 30-Second Take in One Pass

Seedance 2.5 stretches a single generation to 30 seconds with the audio built in, which changes how an ad prompt has to be written. Here is the six-layer structure, the quoted dialogue line that gets lip-sync, and before-and-after rewrites you can paste in.

Mauricio Valdivia

Mauricio Valdivia

·11 min

Seedance 2.5 Prompt Guide for UGC Ads: Direct a 30-Second Take in One Pass

A longer take turns the prompt into a shot list

A media buyer sits down on a Monday with a serum bottle on her desk and thirty seconds of runway to fill. A month ago that was six generations, six prompts, and an afternoon spent hiding the seams between them. Now it is one text box and one take. The text box did not get smarter. It got longer, and a prompt written for a five-second loop, dropped into a thirty-second generation, buys you five seconds of your idea and twenty-five seconds of the model's. This guide covers what changed in Seedance 2.5 that your prompt now has to answer for, the six layers of an ad prompt at this length, the quoted line that produces lip-synced speech, and two full rewrites of the same skincare ad you can adapt to your own product this afternoon.

What 2.5 Changed That Your Prompt Has to Answer For

The 2.0 prompt guide still holds on craft. What it never had to teach was pacing, because at fifteen seconds you could describe one moment well and the model would fill the rest without embarrassing you. That stops being true at double the length.

The take is thirty seconds now

ByteDance's release note is blunt about the headline change: Seedance 2.5 "extends single-pass video generation from 15 to 30 seconds." One pass, one continuous piece of audio and picture, no stitching. For an ad maker that is not a length upgrade so much as a format change. Thirty seconds is the shape of a real spot, which means it wants a beginning, a proof beat, and a close, and a prompt that names only the beginning will get the other two invented for it.

The reference budget moved upstream

The same announcement says users "can now input up to 30 images, 10 video clips, and 10 audio clips as reference materials in a single pass," up from 2.0's nine, three, and three. Read that as a direction, not as your allowance. The same note scopes its own rollout carefully: "Today, Seedance 2.5 is rolling out on Jimeng AI, Doubao Pro, and other platforms, with API access coming soon via BytePlus ModelArk." A vendor ceiling is not the ceiling on the route you are calling, and on Novoads the reference caps stay at nine images, three videos, and three audio clips. The upload contract behind those slots, including how the tags in your prompt actually resolve, is covered in the reference-to-video rules.

What did not change: the instruction surface

Nothing about how you talk to the model was replaced. It still listens for filmmaking vocabulary, it still renders sound described in words, and it still treats a line in double quotes as dialogue. If you already write prompts as a director's brief, none of that work is lost. You are adding one layer, not learning a new language. The folklore that grew up around 2.0, and which parts of it the release actually retired, is unpacked in the workarounds post.

A UGC creator filming a product review without a film crew
Novoads · UGC video ads with AI, ready in minutes.
Try now

The Envelope You Are Writing Into

Novoads runs Seedance 2.5 in production, and the numbers below are the ones the generator and the public API enforce, not the ones a vendor page advertises. Write against these and nothing you ask for gets silently downgraded.

Two resolutions and a full second-by-second grid

The model renders at 480p or 720p. Those are the two arms; there is no higher tier to select. Duration runs the whole integer grid from 4 to 30 seconds, and the API is strict about it: the spec states the model "renders exactly 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 seconds; defaults to 5." A value off the grid is refused rather than rounded, which is the correct behavior and worth knowing before you script to 32 seconds.

Six aspect ratios, one decision

The frame is 9:16, 16:9, 1:1, 4:3, 3:4, or 21:9, and it is a placement decision you make before writing, not after. A 9:16 hook and a 16:9 landing-page loop of the same product are different briefs, because the vertical version has to put the product inside a phone-sized safe area and the widescreen one does not.

Four thousand characters, and why prompts still stay short

The prompt field takes "Up to 4,000 characters for this model." That is roughly six hundred words, and almost no good ad prompt gets near it. Length helps only when it adds a decision to a layer that was empty. A useful discipline: every added sentence should name a thing a camera operator, a wardrobe assistant, or a sound recordist could act on. If it names a feeling instead, cut it.

Audio is part of the same generation. There is no separate voice pass, no surcharge for sound, and no toggle you forgot to switch on. The spoken language follows the words you write: the API notes that "the spoken language comes from the quoted line in your prompt," so a Spanish ad is a Spanish quoted line, not a settings change.

The Six Layers of a 30-Second Ad Prompt

Five of these carried over from shorter takes. The sixth is what the extra fifteen seconds forces you to add.

LayerWhat to writeExample phrase
Subject and shotone person, one action, one framinghandheld vertical close-up
Wardrobe, setting, lightclothing, room, surfaces, named light sourcesgrey tee, tiled bathroom, window light
Product placementwhere the product sits and when it entersbottle already on the ledge, lifted at beat two
Cameramove or lock, stated explicitlyslow push in, then hold
Soundambience, foley, and the quoted linetap water, cap click, one line
Beatsthe order of events across the takethree beats, roughly ten seconds each

The five that carry over

The first five are the layers a good five-second prompt already had. The only difference at thirty seconds is that a missing layer costs you six times as much footage. When a layer is blank the model does not leave a gap, it picks the statistical average, and the average of all advertising is the weightless polished look people scroll past without noticing they did.

The second row deserves more attention than most prompt guides give it, because wardrobe, room and light are what decide whether a clip looks filmed by a person or assembled by a brand. Four things to name:

  • A garment, not a vibe: "oversized grey tee, sleeves pushed up" beats "casual outfit", which the model resolves into catalogue neutral.
  • A room with surfaces: tile, laminate, a chipped windowsill. Surfaces are what the light has to behave against.
  • Light with a source and a direction: "morning window light from camera left" is a direction; "good lighting" is a wish, and the model answers wishes with studio softboxes.
  • One un-styled detail: a hair tie on the wrist, a half-full mug, a phone face-down on the counter. This single item does more for believability than any adjective in the prompt.

The sixth layer: a beat budget

At thirty seconds, write the take as three to five beats and say roughly how long each runs. Beats are not shots you cut between; they are what happens in what order inside one continuous take. Three beats of ten seconds is the safest structure for direct response: hook, proof, close. Then repeat the shared elements across the beats, the same bathroom, the same light, the same hands, or the model treats each beat as a fresh scene and you get three commercials for three different bathrooms.

The quoted line is syntax, not a hint

This is the layer people get wrong most often, and it is documented. fal's Seedance prompting guide states it plainly: "Put any spoken line in double quotes, and the model lip-syncs it, generates the voice, and times it to the cut." Quotes are the instruction. Everything you write around them is scene description, and the model may interpret it, perform it, or print it on screen.

Three practical rules follow from that one sentence:

  • Keep each line under about twelve words. Short lines survive generation cleanly and leave room for the scene's own sound.
  • Give a thirty-second take at most three lines. More than that and the take becomes a monologue, which is the format people scroll past fastest.
  • Put each line inside its beat, written where it happens, rather than stacking all the dialogue at the end of the prompt.

Before and After: One Skincare Ad, Rewritten Twice

Here is the same ad written three ways. The first is what most people type on day one. The second is a five-second hook. The third is the twenty-second version that the longer take makes possible.

The prompt most people type

"A woman shows a skincare product, nice lighting, natural, make it feel authentic and scroll-stopping."

One layer is decided and five are delegated. Read it against the table:

  • Subject and shot: a woman, framed however the model likes.
  • Wardrobe, setting, light: unspecified, so it returns catalogue-neutral clothes in a catalogue-neutral room. "Nice lighting" is not a source or a direction.
  • Product placement: unspecified, so the bottle appears wherever the model puts it.
  • Camera: unspecified, so it drifts.
  • Sound: unspecified, so you get generic room tone or stock music.
  • Beats: unspecified, which at thirty seconds is the expensive one.

The clip comes back competent and anonymous, and the reason it cannot be rescued by adding adjectives is that adjectives were never the missing part.

The five-second rewrite

Prompt 1 · 5-second hook
Handheld vertical close-up of a woman in her late twenties in a small tiled bathroom, oversized grey tee, hair tied up, a half-full mug on the ledge beside her. She lifts a small amber serum bottle into frame and presses one pump onto her fingertips. Slight handheld sway, as if the phone is propped against the mirror. Soft morning window light from camera left plus a warm glow from the mirror bulbs. She says: "Three weeks. That is it." Audio: quiet bathroom room tone, a soft cap click, a fingertip tap on skin.

Every phrase is doing a job, and each one maps back to a layer:

  • Subject and shot: one woman, one action, handheld vertical close-up.
  • Wardrobe, setting, light: grey tee, tied-up hair, tiled bathroom, the half-full mug nobody would have styled in, and two named light sources so the model stops guessing.
  • Product placement: the bottle enters the frame in her hand rather than sitting pre-arranged on a shelf.
  • Camera: the sway is named, and named as imperfect on purpose.
  • Sound: room tone, one foley detail, and six words in double quotes, so they come back lip-synced rather than narrated.

The twenty-second rewrite

Prompt 2 · 20-second three-beat take
A 20 second vertical ad in one continuous take, three beats. Beat one, about six seconds: handheld close-up of a woman in her late twenties at a small tiled bathroom mirror, oversized grey tee, hair tied up. She looks at her skin, sighs, and says: "I gave up on this part of my routine." Beat two, about eight seconds: she lifts a small amber serum bottle from the ledge, presses one pump onto her fingertips and pats it along her cheekbone while the camera pushes in slowly on her hand. She says: "Then I tried one that actually absorbs." Beat three, about six seconds: she steps back, turns her cheek toward the window light, and says: "Three weeks." Keep the same bathroom, the same grey tee, the same morning window light and the same warm mirror bulbs across all three beats. Audio throughout: quiet bathroom room tone, a soft cap click on beat two, a fingertip tap on skin.

Three things changed and only three. The beats are numbered and time-budgeted. The dialogue is split across them instead of delivered in one block. And a single sentence repeats the shared elements, which is the consistency anchor that keeps one take from becoming three. Swap the bathroom for a kitchen, the serum for a supplement, and the foley for a jar lid, and the skeleton holds.

A UGC creator filming a skincare product review on a phone
Novoads · UGC video ads with AI, ready in minutes.
Try now

Four Ways a Take Goes Wrong, and the Layer That Fixes Each

Expect misses. The productive response is surgical: work out which layer failed and change only that one, because rewriting the whole prompt throws away the five layers that worked.

The narrator line the model performs

Write "Voiceover: this serum changed my morning routine" and you will sometimes get exactly that, an actor saying the word voiceover, or a narrator arriving in a clip that was supposed to be a person talking to their phone. The mechanism is documented: fal's guide notes that "The model builds the audio and any on-screen text straight from the prompt." Stage directions are not filtered out on your behalf. Put spoken words in quotes, describe everything else as sound, and never write a label where a line belongs.

Captions you did not order

The same sentence explains the other half of the problem. A phrase that reads like a caption can be rendered as one, burned into the frame at a size and typeface you did not choose. ByteDance says 2.5 "minimizes uncontrolled occurrences in subtitles and background music," and that hedge is the vendor's own word: minimizes, not eliminates, with no published rate underneath it. So the working rule has not changed. Keep on-screen text out of the prompt entirely and add captions afterwards, where you control the font and can localize them without re-rendering.

Too much moving at once

Smeared limbs and objects passing through each other almost always mean the scene has more simultaneous motion than it needs. Count the moving things in your prompt. A converting ad usually needs one, plus the camera. Background crowds, a gesturing second hand, and a camera that pushes while the subject also walks are the three usual culprits, and cutting any one of them tends to fix the frame.

A thirty-second take built on a five-second idea

This is the failure specific to the new length, and it does not look like a bug. The clip is clean, the product is right, and it is dead from second eight onward, because the prompt described one moment and the model looped its energy around it. The fix is the beat budget: if you cannot name three things that happen, generate fifteen seconds instead of thirty. Length is a choice you should have to justify, not a default you take because it is offered.

Before you re-run a miss, walk the layers in order and stop at the first blank one:

  • One subject and one action, or did a second moving thing sneak in?
  • Wardrobe, room, a light source with a direction, and one un-styled detail all named?
  • Is the product's position and entry moment written down?
  • Camera stated as a move or a lock, not left open?
  • Sound described, with spoken words in quotes and nothing else in quotes?
  • Beats numbered and time-budgeted, with the shared elements repeated?

Duration and Resolution Are Prompt Decisions

The two dials change the brief, so pick them before you write. They also change the price, which makes them the cheapest lever you have for testing more angles per budget.

Pick the length before you write

  • 5 seconds, 9:16: one action, payoff in second one. Hook tests live here.
  • 10 to 15 seconds, 9:16: a demo beat plus a visible result. Two beats, one line.
  • 20 to 30 seconds, 9:16: hook, proof, close. Three to five beats, up to three lines.
  • 5 to 10 seconds, 16:9 or 1:1: b-roll with no dialogue, built to survive muted autoplay.

480p is the rehearsal, 720p is the take

Audition prompts at 480p and short lengths, then re-run only the winner at 720p and full length. A prompt that is structurally wrong is wrong at every resolution, so paying the high cell to discover it is a waste. This is the same logic behind creative testing in general: the cheap round exists to kill ideas, and the expensive round exists to finish the survivor.

What a test round costs

From 2026-08-15 the Seedance 2.5 schedule is 16 plus 12 per second in centi-credits at 720p, and 8 plus 6 per second at 480p. In display credits that works out to:

Take length480p720p
4 seconds3.2 credits6.4 credits
5 seconds3.8 credits7.6 credits
15 seconds9.8 credits19.6 credits
30 seconds18.8 credits37.6 credits

A launch promo is currently charging Seedance 2.0's lower rates, so the figure the app shows today is below the schedule above until that window closes on the 15th. Plan against the standing numbers, not the promotional ones.

A realistic first round: five 5-second hook variants at 480p is 19 credits, and one 30-second master at 720p is 37.6. Under 57 credits buys you five tested openings and the finished spot built on whichever one survived. For how the same math looks on the older model, the Seedance 2.0 pricing breakdown has the per-platform comparison, and Kling against Seedance 2.5 covers when a different engine is the better call.

How Novoads Solves the Thirty-Second Brief

Every prompt in this guide runs as written in Novoads. Pick Seedance 2.5, paste the prompt, choose a duration from the 4-to-30-second grid, pick 480p or 720p and one of the six ratios, and the audio comes back inside the same generation. The prompt field takes the full 4,000 characters, so a three-beat brief with wardrobe notes and three quoted lines fits with room left over.

Two product-level details change how you work rather than what you type:

  • Anchor the take to a real photograph. Pass your product shot as the start image and the model animates a real frame instead of guessing at your packaging, which is the reliable answer to invented labels. On this model a start image and a reference set are separate modes, so you pick one per generation rather than sending both.
  • Everything is priced per generation. A five-variant hook round is a decision about credits, not a booking with a studio, which is what makes the cheap-round-then-expensive-round habit affordable in the first place.

Fixing one detail inside a take you already like is a different problem, and it is the one local editing is aimed at.

Access is $1 for 3 days, then $49 a month, cancel anytime. The trial grants 10 credits, which at the schedule above covers two 5-second 480p auditions with change left over: enough to find out whether your product survives contact with a real prompt before you commit to anything.

A UGC creator filming a product review at home
Novoads · UGC video ads with AI, ready in minutes.
Try now

The brief was always the product

Thirty seconds in one pass removes the last technical excuse for a bad ad. There is no stitching to blame, no audio pass that drifted, no editor who did not get the reference. What is left is the thing agencies have always known decides an ad before anyone touches a camera: whether somebody sat down and made the decisions. The shot, the room, the shirt, the moment the product enters, the six words the person actually says, and the order it all happens in. The model will happily make every one of those decisions for you, and it will make them averagely. Write the beats yourself, put the line in quotes, and spend the saved money on testing five openings instead of defending one. There is a clock argument for getting the brief right on the first pass too: measured across 1,045 production renders, a Seedance 2.5 take came back in about four minutes at the median, so a vague prompt costs minutes rather than seconds every time you rerun it. The prompt is the brief, and the brief is still the whole job. Start with one product and one bathroom for $1.

Frequently Asked Questions

What actually changed in Seedance 2.5 for prompt writing?

Length. ByteDance's release note says the model extends single-pass video generation from 15 to 30 seconds. Everything else in a good prompt survives: subject, camera, motion, light, and sound. What a 5-second prompt never had to carry is a time budget, and a 30-second generation given a 5-second idea spends the remaining 25 seconds improvising.

How do I make the actor speak a specific line?

Put the line in double quotes inside the prompt. fal's Seedance prompting guide documents this: a spoken line in double quotes gets lip-synced, voiced, and timed to the cut. It is documented syntax rather than a community trick. Keep lines under about twelve words so they land inside the beat you wrote them into.

Why does my clip have captions I never asked for?

Because on-screen text is generated from the same prompt as everything else. fal's prompting guide states the model builds the audio and any on-screen text straight from the prompt, so a phrase that reads like a caption can be rendered as one. ByteDance says 2.5 minimizes uncontrolled subtitles and background music, but minimizes is the vendor's own word and no rate is published under it. Keep text out of the prompt and burn captions in afterwards.

What length and resolution should I generate at?

On Novoads the grid runs 4 to 30 seconds at 480p or 720p. Use 480p short takes to audition a prompt, then re-run the winner at 720p. A 5-second 9:16 take is the hook test, 15 to 20 seconds fits a demo with a result, and 30 seconds is a full spot with a beginning, a proof beat, and a close.

How much does a Seedance 2.5 clip cost in credits?

From 2026-08-15 the schedule is 16 plus 12 per second in centi-credits at 720p, and 8 plus 6 per second at 480p. That is 7.6 credits for a 5-second 720p take, 3.8 credits at 480p, and 37.6 credits for a 30-second 720p take. A launch promo is charging Seedance 2.0's lower rates until then, so the figure in the app today is cheaper than the standing one.

Can I anchor the prompt to a photo of my real product?

Yes. Pass a start image and the model animates it as the first frame, which is the reliable way to keep a real label sharp instead of letting the model invent one. On Novoads a start image and a reference set are separate modes, so you pick one per generation rather than sending both.

Key Takeaways

  • The change that matters for prompting is length. ByteDance says Seedance 2.5 extends single-pass generation from 15 to 30 seconds, so an ad prompt now has to allocate time, not just describe a moment.
  • A 30-second prompt is a beat sheet. Write three to five beats in order, repeat the shared elements across them, and the model holds one scene instead of re-rolling a new one at every cut.
  • Dialogue has documented syntax. fal's Seedance prompting guide states that a line in double quotes gets lip-synced, voiced, and timed to the cut, so quotes are the instruction and stage directions are noise.
  • Anything you write outside quotes can be performed or printed. The same guide says the model builds audio and on-screen text straight from the prompt, which is why a narrator label ends up spoken and captions arrive uninvited.
  • On Novoads the envelope is 4 to 30 seconds, 480p or 720p, six aspect ratios, and a 4,000-character prompt field, with audio generated in the same call. From 2026-08-15 a 5-second 720p take is 7.6 credits and a 30-second one is 37.6.
Mauricio Valdivia

Mauricio Valdivia

Founder of Novoads

Mauricio is the founder of Novoads, where he works to democratize video advertising with AI for brands in Latin America.