How Long Does AI Video Generation Take? 60 Days of Measured Production Renders
Across 1,045 production renders in 60 days, the median end-to-end wait ran from 45 seconds for a still image to 435 seconds for the slowest video model. Here is the measured table, with the sample size printed beside every number.
Mauricio Valdivia
·11 min

Every Spec but the One You Are Waiting On
You have twenty ad variants to produce and a Thursday deadline, so you want one number: how long does one clip take to come back? The model page has everything else. On fal's page for Seedance 2.5 you can read the token formula, the per-second price at both resolution arms, the aspect ratios, and every field of the input schema down to the seed. On time, the page offers one sentence and no figure: "Long generations take time."
That is not a fal problem, and it is not new. It is what a model page is for. Vendors publish what they charge and what they accept, because those are commitments. Wall-clock time depends on load, queue depth and who else pressed generate in the same minute, so it gets published as advice instead of a number. It is not one page either: ByteDance's own announcement for the same model describes at length what it can do, discusses seconds only as the length of the output, and never says how long producing that output takes. As of 2026-08-10, none of the four vendor pages cited at the end of this post puts a figure on generation time.
So we measured ours. What follows is 1,045 successful generations from Novoads production over the 60 days ending 2026-08-10, across eight models, reported as end-to-end wall time with the sample size printed beside every figure. Nobody else can publish this table, which also means nobody else can check it, so the caveats are in the post rather than in a footnote.
What the Numbers Measure, and What They Do Not
Before any figure is useful, three things about it have to be stated plainly. Skipping them is how a real measurement turns into a false claim.
End to end, not inference
Every figure here is the gap between two timestamps on a production row: when the job was created, and when it was marked succeeded. Four separate things live inside that clock:
- Our own queue wait, before the job is handed to a provider at all.
- The provider's queue, where it sits behind everybody else's renders.
- The render itself, which is the only part a lab benchmark measures.
- Delivery of the finished file, once the pixels exist.
That total is what an operator experiences between pressing generate and having something to watch.
It is emphatically not an inference benchmark. It is not a claim about how fast any model runs on an idle machine, and it is not a measurement of what ByteDance, Google or OpenAI deliver in isolation. If you have seen a lab number for one of these models and it is much lower than what is below, both can be true at once. One of them is measuring the model. The other is measuring the wait.
Ours, on our infrastructure
These are our numbers, from our rows, at whatever concurrency the platform happened to be running at the time. Route the same model through a different provider, a different region, or an empty queue at three in the morning and you get a different distribution.
The value here is not that the figures generalize. It is that they exist, and that they came from real jobs real people were waiting on. Treat them as one operator's honest sample: useful for the shape of the answer and the order of magnitude, not as an industry constant to quote back at a vendor.
Two filters that quietly changed the answer
Two data-hygiene decisions are worth naming, because each one changed the result and neither announced itself.
- The status value is
succeeded, notcompleted. A first pass filtered oncompletedand came back with zero rows, which on a screen looks identical to "there is no data here." A dead sensor and a dead feature are indistinguishable until you go and check the sensor, and that check is what separates a measurement from a conclusion. - A row only counts when its two timestamps are at least 5 seconds apart. Without that floor, rows whose timestamps get written together enter the sample at 0.0 seconds and drag every percentile toward a wait nobody experienced. Applying it moved Seedance 2.0's reported minimum from 0.0 seconds to 120 seconds, a very different story about the fastest render we have ever served.
One more scope line: only successful renders are counted. A job that failed and was rerun appears in this table as its successful run. So this describes renders that worked, not every attempt.

The Table: Eight Models, 1,045 Renders
Here is the whole thing, sorted by sample size. All times are seconds of end-to-end wall clock. Windows are 2026 dates, and the pull is 2026-08-10.
| Model | n | p50 | p90 | Fastest | Slowest | Window |
|---|---|---|---|---|---|---|
| Seedance 2.0 | 792 | 291s | 457s | 120s | 1,682s | Jun 12 to Aug 10 |
| Omni Flash | 69 | 312s | 505s | 71s | 819s | Jul 19 to Aug 6 |
| Seedance 2.0 Mini | 61 | 133s | 183s | 77s | 274s | Jul 8 to Aug 6 |
| Veo 3.1 | 34 | 122s | 208s | 65s | 334s | Jun 12 to Aug 8 |
| Seedance 2.5 | 34 | 234s | 326s | 85s | 402s | Aug 7 to Aug 10 |
| Nano Banana Pro (image) | 21 | 45s | 62s | 31s | 180s | Jun 30 to Jul 13 |
| Sora 2 | 17 | 155s | 207s | 75s | 278s | Jun 13 to Aug 9 |
| Sora 2 Pro | 17 | 435s | 830s | 283s | 2,177s | Jun 12 to Aug 5 |
Read the n column first
Sample size is doing as much work in this table as the percentiles are, and the rows are not equally trustworthy.
- Seedance 2.0 rests on 792 renders. Its p50 and p90 are stable enough to plan a week against.
- Omni Flash (n=69) and Seedance 2.0 Mini (n=61) are solid, if narrower.
- Veo 3.1, Seedance 2.5, Sora 2 and Sora 2 Pro rest on 17 to 34 renders each. Read their p50 as a rough centre and their p90 as a weak estimate. A 90th percentile computed over seventeen observations is one unlucky render away from moving, and it deserves none of the confidence the Seedance 2.0 row has earned.
Models with fewer than 15 successful renders in the window are not in the table at all. Showing them would mean publishing a percentile no sample supports, which is the failure mode this whole post exists to avoid.
What the window column is doing
The window is when the renders actually happened, and it is not the same for every row.
- Seedance 2.5 is four days old. All 34 of its renders landed between August 7 and August 10, during the model's rollout. That is not a settled steady-state figure, and anyone quoting 234 seconds as its number should quote the four days alongside it.
- Nano Banana Pro is the stalest row. Its most recent render in this sample is from July 13, four weeks before the pull.
- Omni Flash stops on August 6, so its 69 renders say nothing about the four days after that.
A number is only as current as the window it came from, which is why that column sits in the table rather than in a footnote. If you want the model itself rather than its clock, we covered what Seedance 2.5 is separately.
Output Length Is Not Render Time
The most common planning mistake is to reason from clip length: an eight-second clip feels like it should be quicker than a fifteen-second one, and a ten-second clip feels like it should not take long at all. Inside one model that intuition mostly holds. Across models it collapses entirely.
A clip measured in seconds, a wait measured in minutes
Google's documentation says Veo 3.1 produces native clips of 4, 6 or 8 seconds, and DeepMind's own model page states flatly that "Veo videos are 8 seconds long." Our median wait for one of those is 122 seconds (n=34).
That is the whole lesson in one line. The output is measured in seconds and the wait is measured in minutes, and there is no ratio you can carry from one model to the next. Veo 3.1 has the fastest median in our table and a short native clip. Sora 2 Pro has the slowest median. You cannot rank these models on speed by looking at how long their videos are.
The clearest case is the one with a fixed length
Omni Flash is where the arithmetic is unambiguous, because Google stated in June that "Omni offers 10-second video generations currently, with longer durations coming soon." One length, no dial to argue about.
Our median for it is 312 seconds (n=69), roughly thirty times the length of the clip that comes back, and its p90 is 505 seconds. If you have only ever reasoned about this model from what Omni Flash does, thirty times is the number that changes how you schedule it.
Stills are a different order of magnitude
One row in the table is not a video model at all. Nano Banana Pro generates images, and it sits here purely as a scale marker. It is never averaged into any video figure in this post, and it should not be averaged into one anywhere else.
Its median is 45 seconds and its p90 is 62 seconds (n=21). A still comes back in about a minute; the video models on the same board sit between two and seven and a half minutes at the median. That gap is the practical argument for settling the frame first and animating second, which is how product videos built from a photo tend to get made in the first place.

Plan on the p90, Not the p50
A median is the number you quote in a meeting. A 90th percentile is the number that decides whether the batch is finished before you leave. Once you are producing more than one clip, only one of those two is a planning tool.
The median is the number you remember
Half of your renders come back faster than the p50. One in ten takes longer than the p90. On a batch of twenty that means you will cross the p90 about twice, by construction, not by bad luck. Planning the batch on the median therefore guarantees you are surprised, and the surprise always lands at the end of the queue when there is no time left to absorb it.
A worked example: twenty variants
Take twenty Seedance 2.0 variants and run them one after another.
- At the p50 of 291 seconds, that is 5,820 seconds, about 1 hour 37 minutes.
- At the p90 of 457 seconds, it is 9,140 seconds, about 2 hours 32 minutes.
The same batch of twenty has a 55-minute spread between the optimistic read and the cautious one. Book the afternoon on the median and the last few variants land after you have gone home.
Now change the model instead of the plan. Twenty on Seedance 2.0 Mini takes about 44 minutes at its p50 of 133 seconds, and about 1 hour 1 minute at its p90 of 183 seconds. Put the two side by side and the interesting line falls out on its own: the cautious number for the smaller model is still faster than the optimistic number for the larger one. Model choice moves your schedule further than scheduling discipline does.
The tail is where the schedule breaks
The slowest column is not decoration. Seedance 2.0's slowest successful render in 60 days took 1,682 seconds, about 28 minutes, against a median of 291. Sora 2 Pro's slowest took 2,177 seconds, about 36 minutes, on a p50 of 435 and a p90 of 830 (n=17).
Those are single observations, not rates, and that is exactly how to use them. They are not something to plan around. They are the answer to "is this render stuck or is it just slow", which is a question worth being able to answer without cancelling a job that was going to finish.
Why One Model Ranges From Two Minutes to Twenty-Eight
A spread of 120 to 1,682 seconds inside one model is not noise, and it is not a sign that anything is broken. It is what an end-to-end measurement looks like when it is honest about everything it contains.
Queue depth is part of the clock
Because the measurement runs end to end, everything between creating the job and the finished file existing is inside it, including two queues: ours and the provider's. A render submitted into a busy minute waits behind other people's renders before a GPU ever looks at it.
That wait is completely invisible from the outside, it is absent from every vendor page, and it is a real part of what you experience. It is also the single biggest reason a lab benchmark and an operator's stopwatch disagree.
The dials a real job used move the number
Within a model, settings move the clock, and vendors say so without quantifying it. fal's page for Seedance 2.5 gives the direction and no figure: its guidance is to "use the queue API rather than a synchronous call for anything past a few seconds of output." Three dials on a real job all add work:
- Duration, which is what the vendor's own warning is about.
- Resolution, since the work scales with the frame area being generated.
- Generated audio, produced alongside the picture rather than layered on afterwards.
The ceiling has been moving too. ByteDance's Seedance 2.0 supports "15-second high-quality multi-shot audio-video output" in a single pass, and its 2.5 announcement says "Seedance 2.5 extends single-pass video generation from 15 to 30 seconds." A 30-second single take is a better ad unit than six stitched clips, and it is also more compute in one job. Longer output does not arrive without a longer clock, which is worth remembering before reading the four-day Seedance 2.5 row as a promise about the future.
What the table cannot see
Three limits, stated so nobody reads more into the data than it holds.
- Failed renders are absent. Only successful jobs are counted, so this is the shape of renders that worked.
- Per-render settings are not broken out. Each row mixes whatever durations, resolutions and audio settings real production jobs used.
- Nobody else's infrastructure is in here. Every figure is ours, which is the point and also the limit.
What This Changes About Planning a Batch
Knowing one render takes 291 seconds is trivia. Knowing what that does to a week of creative testing is the reason to measure it.

Plan iterations per hour, not seconds per clip
The useful unit is not how long one render takes. It is how many rounds of creative you get through in an afternoon, which is the real constraint on any serious ad creative testing program.
Run sequentially at the p50, one hour buys roughly:
- 29 Veo 3.1 renders
- 27 Seedance 2.0 Mini renders
- 15 Seedance 2.5 renders, on four days of data
- 12 Seedance 2.0 renders
- 8 Sora 2 Pro renders
At the p90 the Seedance 2.0 figure falls to about 8 an hour and Sora 2 Pro to about 4. That is the range a testing plan actually lives in, and it is why the p90 matters more than the p50 the moment you are planning a batch rather than a single clip. If you are still sizing how many variants a test needs, how many ad creatives you need is the other half of this arithmetic.
Since each attempt costs minutes rather than seconds, the prompt is where the time is won or lost: directing a 30-second Seedance 2.5 take in one pass is cheaper than discovering the framing was wrong on the third render. The pressure is sharpest on channels that report at the ad level and no deeper, which is exactly the position advertisers running ChatGPT ads are in right now.
Match the model to the clock, not only to the look
The fast model and the good model are not always the same model, and they do not have to carry the same job.
Reach for the quicker end of the table when:
- You are testing angles, hooks or script structure rather than finish.
- You need twenty of something so you can throw away seventeen.
- The batch has to be reviewable in one sitting.
Reach for the heavier end when:
- One angle has already proven itself and is going live.
- The clip is the hero asset rather than a probe.
- You are rendering once, not twenty times.
Moving a single winner up to a heavier model costs one render's wait instead of twenty, which makes this a scheduling decision as much as an aesthetic one. It is the part of creative operations that model comparison posts tend to skip. When you do compare the heavier engines head to head, Seedance against Veo for ads is the pairing that matters most here, because those two sit at opposite ends of our measured clock.
Buffer the tail, then stop watching
Add the p90 to the plan, add a little on top of it for the tail, and then leave the batch alone. The one behaviour worth unlearning is the reflex to cancel and resubmit a render that has passed the median, because on the slower models in this table a job sitting at eight minutes is inside the ordinary range rather than broken. Cancelling it restarts the wait from zero and puts you further behind than doing nothing would have.
How Novoads Closes the Render-Time Blind Spot
Every model in the table above is one we run in Novoads in production, which is the only reason this table can exist: the clock is measured on our own rows rather than estimated from a vendor page. Choosing the quick model for a first pass and the heavier one for the winner is a choice inside one tool rather than a migration between two, so the scheduling logic in this post is something you can apply in an afternoon rather than a quarter. You can start a project and watch your own clock instead of inheriting ours.
The Clock Is the Real Creative Budget
Teams plan creative testing around what they can produce. The binding constraint is usually simpler and much less discussed: how many finished renders exist by Friday. That number is set by a distribution nobody publishes, which is how a team ends up planning a week around a figure it has never actually seen.
A model's speed is not a spec you can look up. It is a distribution you have to measure, and the tenth-slowest render out of every hundred is the one that decides whether the campaign ships on time.
Frequently Asked Questions
How long does AI video generation actually take?
Measured end to end on 1,045 successful production renders over the 60 days ending 2026-08-10, median waits were 122 seconds for Veo 3.1, 133 seconds for Seedance 2.0 Mini, 155 seconds for Sora 2, 234 seconds for Seedance 2.5, 291 seconds for Seedance 2.0, 312 seconds for Omni Flash and 435 seconds for Sora 2 Pro. A still image from Nano Banana Pro came back in 45 seconds at the median. These are wall-clock waits on our own infrastructure, not vendor inference benchmarks.
Why do model pages not publish a generation time?
Because a render time is not contractual and price is. A vendor page documents what it charges, what resolutions it offers and what fields the API accepts, since those are commitments. Wall-clock time depends on load, queue depth and who else is generating at that moment, so it is published as guidance rather than a number. fal's Seedance 2.5 page is a good example: it prices per second of generated video and says only that long generations take time.
Does a longer clip take longer to generate?
Inside a single model, generally yes, and vendors say so without quantifying it. Across models the relationship breaks completely. Google states that Omni Flash currently produces 10-second clips, and our median wait for one is 312 seconds, roughly thirty times the clip length. Veo 3.1 generates native clips of 4, 6 or 8 seconds, and our median wait for one is 122 seconds. Output length and render time are two different axes.
Should I plan against the median or the p90?
The p90, once you are producing more than one clip. Half of your renders finish faster than the p50 and one in ten takes longer than the p90, so a batch of twenty will cross the p90 about twice by construction. Twenty Seedance 2.0 variants run sequentially take about 1 hour 37 minutes at the p50 and about 2 hours 32 minutes at the p90. That 55-minute spread is the difference between a batch that finishes before you leave and one that does not.
Why do the same model's renders range from 120 to 1,682 seconds?
Because the measurement is end to end, everything between creating the job and the finished file existing sits inside it, including our queue and the provider's queue. On top of that, the settings a real job used move the clock: duration, resolution and audio generation are all more work. The table mixes whatever settings production jobs actually ran with, which is exactly why the spread is wide and why the maximum is worth knowing.
Do these numbers apply to other platforms running the same models?
No, and they should not be read that way. Every figure is measured on Novoads production rows, on our infrastructure, at whatever concurrency the platform was running at the time. A different provider, a different region or an emptier queue produces a different distribution. What transfers is the shape of the answer: waits measured in minutes for clips measured in seconds, wide tails, and a p90 that decides your schedule.
Key Takeaways
- Vendor model pages document price, resolution and the full input schema, and state no generation time. fal's Seedance 2.5 page says only that long generations take time, with no number attached.
- Measured over 60 days on 1,045 successful Novoads production renders, median end-to-end waits ran from 122 seconds (Veo 3.1) to 435 seconds (Sora 2 Pro), with a still image coming back in 45 seconds.
- These are end-to-end waits, not inference benchmarks: the clock includes our queue, the provider's queue, the render and the delivery. They are our numbers on our infrastructure at our concurrency.
- Sample size decides how much to trust each row. Seedance 2.0 rests on 792 renders; Veo 3.1, Seedance 2.5, Sora 2 and Sora 2 Pro rest on 17 to 34, so their p90 is a weak estimate.
- For planning a batch, the p90 matters more than the p50: twenty Seedance 2.0 variants run back to back take about 1 hour 37 minutes at the median and about 2 hours 32 minutes at the p90.




