AI Video Physics: Seedance 2.5, Veo 3.1 and Kling 3.0 Pro Failed a Water Pour
In a September 23, 2026 write-up, HumanSignal gave Seedance 2.5, Veo 3.1 and Kling 3.0 Pro the same first frame of a water pour, and none produced a convincing one. Here is which product-ad shots to plan around, and the clay-previs workaround that saved their spot.
Mauricio Valdivia
·11 min

The character held. The water did not.
HumanSignal wanted a simple ad of about 20 seconds. A friendly household robot waters a houseplant, the camera swings around, and the viewer discovers the stream is missing the pot entirely. It did not stay simple.
In a post dated September 23, 2026, the company's creative director, Todd Morey, explains what held and what broke. Once the team fed Seedance 2.5 a character turnaround and two views of the room, the robot kept the same face and proportions from shot to shot. The near-miss gag and the liquid were another story. So the team ran an informal side-by-side: Seedance 2.5, Veo 3.1 and Kling 3.0 Pro each got the same first frame of a real water pour, and by HumanSignal's account all three failed to give a convincing account of pouring water, each in its own way.
Novoads runs all three of those models, which is why this write-up matters to anyone making product and beverage ads. Below: what HumanSignal tested, what its results do and do not show, which shots to plan around, and the workaround that saved the spot. Novoads did not re-run any of these tests. Every result in this post is HumanSignal's.
The ad that became a stress test
A near-miss is the whole joke
HumanSignal builds datasets, including novel data for physical AI, and it wanted to test whether current video models could carry a real ad campaign. The team storyboarded an amiable household robot that seems to succeed until the camera moves and reveals the task as almost, but not quite, done. The post describes two of the spots:
- Ad 1: the robot waters a houseplant, the camera pivots, and the stream misses the pot entirely.
- Ad 2: the robot delivers a tray of coffees and walks it straight into a glass door.
The surprise shift from success to failure was meant to be the payoff and the humor. According to the post, the camera orbit, the task failure and the fluid dynamics were all extremely challenging to pull off with the current generation of video models.
Three hard problems inside one shot
In hindsight, the post calls the little ad a triple stress test:
- A novel task outcome. A stream that almost reaches the plant is an unusual result for a watering scene.
- A camera pivot through a 3D scene. The reveal only works if the room stays coherent as the viewpoint moves.
- Fluid dynamics. The post calls it notoriously hard.
Our read: most product ads need only one of those at a time. A beverage spot rarely needs a deliberate failure, but it almost always wants a pour. That is why the pour test is the part of the post to study.
Read it as practitioner notes, not a verdict
Two cautions before the results:
- The author has a stake. Novel datasets for AI training are what HumanSignal does, and the post closes with a pitch for exactly that, so a finding that video models misread physics suits its business.
- The post labels its own evidence. These examples, it says, are not a field-wide measure of progress.
Treat what follows as one team's saved examples on three specific model versions, not a benchmark and not a verdict on AI video in general.

What held: one robot, shot after shot
The character turnaround
Going in, the team's biggest worry was character consistency, a robot that mutates between shots. That turned out to be the manageable part. The fix was a character turnaround: multiple views of the robot, fed to the model as an image reference. The post shows four views of the wheeled robot used in the ad.
With the turnaround in place, HumanSignal writes that Seedance 2.5 held the same character across the spot admirably well. Same face, same proportions, same joint design, shot after shot.
Two views of the room for a moving camera
To keep the set stable while the camera moved, the team also generated a three-quarter view of the room, with one corner showing how the key objects relate to each other, plus a straight-on view. Both went in as references. What the team expected to be the hard problem, in its own words, turned out to be the least of its worries.
The product-ad version of a turnaround
Our reading: the trick maps directly onto products.
- Product turnaround: the front, the side, the back and a tight shot of the label.
- Set reference: the counter, desk or bathroom shelf the ad lives on, shot from two angles.
- One job per reference: the product views govern how the product looks, the set views govern where things sit.
Our advice: if a bottle changes shape between cuts, fix the references before you blame the model. That is the move that held HumanSignal's robot steady. Our guide to keeping one AI actor consistent across ad variations covers the same discipline for people, and the Seedance 2.5 prompt guide covers how that model wants its prompts written.
The pour: three models, three different failures
How the side-by-side worked
The setup, as the post describes it:
- Ground truth: a real video of water being poured, a bit clumsily by the team's own admission, from a full pint glass into an empty one.
- Input: the first frame of that clip, sent as an identical image to every model.
- Prompt: an identical action-only prompt for Seedance 2.5 (ByteDance), Veo 3.1 (Google DeepMind) and Kling 3.0 Pro (Kuaishou).
- Task: pour about 60% of the water across.
The prompts described the steps, not their physical consequences. The post's example is "pour the water," not "the water level falls," which leaves the consequences for each model to infer. One caveat on wording: the task line shown on the page is marked as summarized from the experiment notes and not a verbatim prompt log. It is a summary of the instruction, not the exact prompt.
What came back
HumanSignal's verdict is blunt: "All three failed to give us a convincing account of pouring water, and they failed in different ways." The shared weakness was volume transfer. The visible levels in the two glasses do not change the way they should.
Kling 3.0 Pro gave the sharpest illustration. To the team's eyes, its pour was the most realistic of the three, but the end result was essentially two glasses of water, as if the scene had created water from nowhere. The page's viewer tells readers to follow the water level in both glasses, especially after the pour stops.
A worked check for any pour shot
The 60% task doubles as a check you can run on your own generations. For two matching glasses:
- Glass A starts full (100%) and should end near 40%.
- Glass B starts empty (0%) and should end near 60%.
- The total should read one glass at every frame, including after the pour stops.
Two tells to scrub for: a total above one glass, which is the water-from-nowhere result HumanSignal describes, or a source glass that barely drops while the other fills. It takes a few seconds on the timeline, and it catches the one error a drinks brand cannot run.
Looking right is not behaving right
The post draws a distinction worth keeping in mind whenever a clip looks great:
- Realism: textures, lighting, motion blur, the surface statistics of the world.
- World understanding: persistent objects, conserved quantities, consistent geometry.
In the post's words, a model can produce convincing pixels while quietly duplicating water or bending a joint that doesn't exist.
HumanSignal points to the Physics-IQ paper, which found that visual realism and physical understanding were largely unrelated in the models it tested, and to VideoPhy-2, which also found shortcomings in physical commonsense, including conservation laws. The post adds that those studies evaluated their own model sets, not the three versions it tested.
The robot arm and the shirt: a more mixed picture
The toy arm got re-engineered mid-shot
The second test used an image of a Radio Shack Armatron, the 1980s toy robot arm, with a blue ball and a plastic cup on a table. Each model was asked to have the arm pick up the ball and drop it in the cup. The team chose the Armatron because much of its mechanism is legible from a single still.
According to HumanSignal, the models improvised the machine's kinematics instead:
- Geometry drift: wrist and arm geometry warped mid-motion to reach the cup.
- A hop: two of the models had the whole toy hop closer to the cup.
- Right goal, wrong machine: the outputs worked toward the goal, but the mechanism changed along the way.
The post's summary is the line worth remembering: the models appeared to understand the goal, and they hallucinated the machine.
The shirt fold got better
The balance matters here. HumanSignal also re-ran a robot-arm shirt-folding task from an earlier 2025 Google Veo failure example, using the same starting frame against the current models. Two of the three, Seedance 2.5 and Veo 3.1, now produce a more credible fold, with the cloth's movement more closely connected to the arms' actions. In the earlier clip, the post says, the transition to a folded state was much less convincing. In the same sentence, the team says these examples are not a field-wide measure of progress, but are an encouraging practical improvement.
What the three tests add up to
- Pour: all three unconvincing, each failing differently, per HumanSignal.
- Toy arm: goal understood, mechanism warped along the way.
- Shirt fold: Seedance 2.5 and Veo 3.1 more credible than the 2025 example.
HumanSignal's own summary is that across these saved examples, it kept finding a gap between plausible appearance and consistent behavior. Our reading is narrower than a headline would like. The gap showed up where a quantity had to be conserved or a mechanism had to keep its shape, and at least one task, the fold, has visibly improved since 2025. For the specs that do separate these engines, see Kling vs Seedance for AI video ads and Seedance 2.0 vs Veo 3.1 for video ads.

Which product-ad shots to plan around
Everything in this section is our reading of HumanSignal's results for product and beverage advertisers. The post did not test these ad shots.
Pours, fills and anything that changes volume
A drink poured over ice, a sauce drizzled on food, a serum dropper filling, protein powder scooped into a shaker. Each asks the model to move a quantity from one place to another and keep the total honest, the exact job all three models fumbled in HumanSignal's pour test. Plan these so the volume change is either real or never on screen:
- Film the pour for real. A phone on a tripod and one take of the actual product gives you the one shot the models fumbled.
- Cut around it. Generate the full bottle and the filled glass as two separate shots, and let the edit imply the pour.
- Keep it tight. A shot that cuts as the stream starts leaves the model less volume to account for than a wide shot that lingers after the pour stops, the moment HumanSignal tells viewers to watch.
Mechanical demos: hinges, pumps, folding parts
The Armatron result is the warning for any product whose selling point is how it moves: a stroller that folds, a pump bottle, a blender lid that locks, a phone stand that pivots. HumanSignal's models kept the goal and changed the machine. A demo that shows a hinge your product does not have is worse than unconvincing: it misrepresents the product.
If the mechanism is the claim, show it in real footage or a clean still. Let the AI shots carry what they handle well: a person holding the product, reacting, talking to camera. Our guide to making product videos with AI walks through that split from one product photo.
Deliberate failure gags
HumanSignal's internal name for this is failure choreography: prompting for a deliberate, plausible-looking failure. Its working explanation is what it calls a competence prior. Show a model a robot, a pitcher and a plant, and it tends toward the water landing where water should land. The team presents that as its working explanation, not a proven mechanism.
For ad makers, the practical point stands either way. A spilled drink, a missed catch or a comic "before" disaster is the same kind of shot as HumanSignal's near-miss, the part of its ad that, per the post's own heading, failed completely. If the joke depends on the miss, get the miss from real footage or from previs, the workaround below.
A shot list for a 15-second beverage ad
Here is how we would split a 15-second iced-coffee ad after reading the post. The AI shots keep every liquid level still; the two shots that need real physics are filmed.
| Shot | Seconds | Source | Why |
|---|---|---|---|
| Actor holds the can | 0-3 | AI video | No liquid moves |
| Ice drops into glass | 3-5 | Real footage | Physics on screen |
| The pour | 5-8 | Real footage | Volume must add up |
| Actor talks, glass in hand | 8-12 | AI video | Level stays still |
| Can beside full glass | 12-15 | AI video | Static packshot |
Two real shots of five, filmed on a phone in one sitting, and the rest generated. If you have not broken a script into shots before, our walkthrough on how to turn a 30-second script into a shot list is the place to start.
The workaround that saved the spot: clay previs
Borrow the geometry from a 3D tool
The practice that saved HumanSignal's robot ad was to borrow spatial structure from a system that represents it explicitly. In order:
- Build a clay previs. The team used OpenAI's Astra, which the post elsewhere calls GPT-6 Astra, to help generate a clay previsualization of the action and camera move.
- Keep it crude. Untextured geometry stood in for the plant, the pitcher and the robot.
- Feed it with one job. The previs clip went to the video model as a tagged reference governing motion and camera only.
On the page, the clay previs is a 6.5-second camera orbit, and the finished ad's current cut runs 19.9 seconds.
Less detail was better
The counterintuitive part, in HumanSignal's words: less detail was better. Blocky, simple, deliberately unfinished geometry was less likely to bleed its visual style into the final output. The team wanted the model to take the choreography and ignore the clay.
The post calls the result a practical substitute for the spatial consistency it could not reliably get from the video model alone. For now, it says, previs is an essential part of the work, especially for novel scenarios where the model cannot be relied on to have seen the same thing before.
Tagged references do the steering
The previs only helps because the model can be told what a reference is for. HumanSignal writes that Seedance 2.5 supports up to 50 image, sound and video references, and that each can be tagged within a prompt and given a specific job, such as governing geometry and camera placement only. ByteDance's own launch post puts the ceiling at up to 30 images, 10 video clips and 10 audio clips as reference materials in a single pass, which adds up to HumanSignal's 50.
The same ByteDance post says the model strengthens a range of reference capabilities, including clay render, motion and creative references. That is the maker naming the same kind of clay input HumanSignal leaned on. How many reference slots a given app exposes is that app's choice, so check the upload screen before you plan a shot around dozens of them. For the words that steer a camera move without any reference at all, see our list of camera movement prompts for ads.
How Novoads fits a physics-heavy brief
Novoads runs all three engines from HumanSignal's side-by-side: Seedance 2.5, Google Veo 3.1 and Kling v3 Pro, the Kling 3.0 Pro tier. Our reading of the post is simple to act on in Novoads:
- Generate the shots where nothing has to be conserved: the actor, the hold, the packshot.
- Film the few where a level or a mechanism has to add up.
- Compare a risky shot on more than one engine before you commit the edit to it.

Retries are where the budget goes. At the Plus plan's credit rate, a clip costs from about 25 cents for a five-second Seedance 2.0 Mini clip to about $4 for an eight-second Google Veo 3.1 clip, and about $10 for a full 30-second Seedance 2.5 take. A pour that needs five attempts is five paid clips, while a real pour on a phone costs a glass of water. You can start a project on Novoads and see every plan on the pricing page.
Plan for the physics, generate the rest
HumanSignal's post is not proof that AI video cannot handle liquids. It is one team's honest notes on three model versions, with a failure, a partial success and a fix. The fix is the lesson: give the model structure it cannot invent, and keep the shots that need real physics real. The models can hold a face. Until they can hold a glass of water, film the pour.
Frequently Asked Questions
Did AI video models fail HumanSignal's water pour test?
The three models HumanSignal tried did. In a post dated September 23, 2026, the company says it gave Seedance 2.5, Veo 3.1 and Kling 3.0 Pro the same first frame of a real pour from a full pint glass into an empty one and asked each to move about 60% of the water across. Its verdict: all three failed to give a convincing account of pouring water, and they failed in different ways. The common struggle was volume transfer. The test covered those three model versions only, so it says nothing about other models.
Which models did HumanSignal compare?
Seedance 2.5 from ByteDance, Veo 3.1 from Google DeepMind and Kling 3.0 Pro from Kuaishou. Each got an identical first-frame image and an identical action-only prompt, across three tasks: pouring water between two glasses, a toy robot arm dropping a ball in a cup, and a robot-arm shirt fold. Novoads did not re-run these tests; every result is HumanSignal's.
Is this a benchmark of AI video physics?
No. HumanSignal describes a simple experiment and calls the clips saved examples, and it says plainly that these examples are not a field-wide measure of progress. The task descriptions on the page are marked as summarized from experiment notes, not a verbatim prompt log. Read it as one production team's informal side-by-side, useful for planning shots, not for ranking models.
What is clay previsualization in AI video?
In HumanSignal's workflow, it is a rough 3D mock-up of the action and camera move, made of crude, untextured geometry standing in for the real objects. The team fed that previs clip to the video model as a tagged reference governing motion and camera only. The post says less detail was better, because blocky geometry was less likely to bleed its look into the final frames.
How many references does Seedance 2.5 accept?
ByteDance's own Seedance 2.5 launch post says users can input up to 30 images, 10 video clips and 10 audio clips as reference materials in a single pass. HumanSignal's post counts that as up to 50 references and adds that each can be tagged in the prompt with a specific job. How many of those slots a given app exposes depends on the app, so check the upload screen you are working in.
Can I use these three models in Novoads?
Yes. Novoads runs Seedance 2.5, Google Veo 3.1 and Kling v3 Pro, the Kling 3.0 Pro tier. The practical read of HumanSignal's post is to plan around shots where a liquid level or a mechanism has to stay consistent, film those for real, and generate the rest. Plans are listed on the Novoads pricing page.
Key Takeaways
- HumanSignal's September 23, 2026 post reports that Seedance 2.5, Veo 3.1 and Kling 3.0 Pro, each given the same first frame, all failed to give a convincing account of pouring about 60% of a full pint glass into an empty one. The common struggle was volume transfer.
- It is one team's saved examples, not a benchmark. The post says its examples are not a field-wide measure of progress, and its task lines are summaries from experiment notes, not verbatim prompts.
- Not everything failed. A character turnaround plus two room views kept the robot consistent in Seedance 2.5, and Seedance 2.5 and Veo 3.1 folded a shirt more credibly than a 2025 Veo failure example.
- For product and beverage ads, plan around pours, fills, mechanical demos and deliberate failure gags: film those for real or cut around them, and let AI carry the shots where nothing has to be conserved.
- The workaround that saved HumanSignal's ad was a crude clay previs fed as a tagged reference for motion and camera only. ByteDance's own launch post says Seedance 2.5 strengthens clay render references.




