H3 Max Lip Sync on fal: a 15-second talking-head ad costs $0.75 to $4.80
fal made H3 Max Lip Sync, its post-train of MiniMax's open-weight H3, public on September 18, 2026: it turns one still image and a 5 to 15 second voice track into a lip-synced clip billed at $0.05 to $0.32 per second, which prices a talking-head ad and caps every call at 15 seconds.
Mauricio Valdivia
·12 min

The voiceover ran 22 seconds. The clip stopped at 15.
A skincare founder pairs her best headshot with a voiceover she recorded on her phone and sends both to fal's newest talking-head endpoint. The clip comes back convincing. It is also short. The last seven seconds of her script, the part where she names the offer and asks people to tap, never made it into the video.
Nothing broke. H3 Max Lip Sync, which fal made public on September 18, 2026, generates a video from an image and supplied audio, and its schema describes what happens to long audio in one flat sentence: audio over 15 seconds is automatically clipped to its first 15 seconds. The endpoint bills per second of generated video, from $0.05 at 480p to $0.32 at 2K, so a full 15-second clip runs $0.75 to $4.80 before anyone has been paid for the voice.
That is a useful price for a talking head, attached to a shape you have to build around. Below: what went live, what each input decides for an ad, what a clip and a finished spot cost, and where the endpoint sits next to the talking-actor route in Novoads, which does not run H3.
What fal put live on September 18
The model page carries no date. fal's catalog does, and it answers the two questions that matter for a news story: what exactly shipped, and who built it.
One still and one voice track in, one talking clip out
fal's own summary of the endpoint is two sentences: H3 Max Lip Sync generates a video from an image and supplied audio, synchronizing mouth movements to the soundtrack. It supports optional transcription guidance and output resolutions from 480p to 2K.
The API schema is just as short. Two inputs are required, image_url and audio_url, and the response returns three fields: the video, the seed that produced it, and the output duration in seconds. fal describes that video as lip-synced video with the supplied soundtrack, which is the sentence that matters most to an advertiser. What the viewer hears is exactly what you uploaded. The script, the voice, the language and the pacing are all settled before the call.
In fal's catalog the entry was created on September 17 at 08:31 UTC and switched to "status":"public" with "publishedAt":"2026-09-18T05:26:34.782Z". That timestamp is the event this post is about: the moment anyone holding a fal key could call it.
The eighth H3 Max endpoint in 23 days
Query fal's catalog for H3 Max and it returns eight entries. Lined up by their publication stamps, they read like a release train:
- August 26: text-to-video and image-to-video.
- August 31: reference-to-video.
- September 2: two Turbo variants, one for text and one for images.
- September 3: Director, the realtime stream you steer mid-take, which we priced in H3 Max Director's session billing.
- September 11: Camera Controls.
- September 18: Lip Sync.
The newest arm is also the narrowest. Its input schema has no prompt field at all. It does not invent a scene. It takes one face and one recording and makes the face say the recording, which is the shot every talking-head ad is built on.
fal's post-train, filed under MiniMax's name
The endpoint path starts with minimax/, and fal's catalog files the entry under "modelLab":"Minimax" and "modelFamily":"H3". Every label on it points at MiniMax. The provenance sentence on fal's own H3 Max page points somewhere else: H3 Max is post-trained by fal on top of the open-weight base MiniMax H3 model. The same page adds that standard MiniMax H3 is a separate frontier model with its own endpoints.
So the credit splits cleanly. MiniMax published the open weights. fal post-trained them and built the lip-sync endpoint on top. If you want the base model and its native-audio claims, our explainer on what MiniMax H3 is covers that side. This post stays on fal's lip-sync arm and what it does for a talking-head ad.

The input contract, read as an ad brief
The schema has six input fields. Each one is a decision somebody used to make on set, and reading them as a brief tells you what the endpoint will and will not decide for you.
The still is the casting
The image rule is one line: the aspect ratio must be between 0.4 and 2.5. Every standard ad placement fits inside that window:
- 9:16 vertical is 0.56.
- 4:5 feed is 0.8.
- 1:1 square is 1.0.
- 16:9 landscape is 1.78.
The resolution field adds a second rule worth reading twice: output uses the supported aspect ratio nearest the image. The endpoint picks the frame, based on your photo. So crop the still to the placement you are buying before you upload it, measured against the spec each platform actually wants, rather than letting a near-miss ratio choose your framing. fal's playground accepts jpg, jpeg, png, webp, gif, avif, heic and heif.
Everything else about the performer is fixed by that one photo: the face, the resting expression, the wardrobe, the room, and whatever is in their hands. If the product should be on camera, it has to be in the still. The same logic runs through every AI avatar made from a photo, and so does the obligation that comes with it: animate a face only when the person in it agreed to be animated.
The audio is the script
The audio field carries three rules in one description. Audio must be at least 5 seconds long. Audio over 15 seconds is automatically clipped to its first 15 seconds. And the output video matches the clipped audio duration in a single generation.
Read as a brief, that means:
- The shortest clip you can buy is five seconds. A two-second hook sting cannot be generated on its own; it rides inside a longer line.
- The longest clip per call is fifteen seconds. Anything longer is two calls or a cut.
- The clip is exactly as long as the audio. There is no separate duration setting to fight with.
The playground takes mp3, ogg, wav, m4a and aac. What goes into that file is entirely yours: a founder recording, a hired voice actor, or a synthetic voice from whichever text-to-speech tool you already pay for. Delivery, accent and pace are decided in the booth, not by the model.
Two switches and a seed
The remaining four fields are small, and two of them are on by default.
enable_transcription, on by default. fal's description: transcribe the supplied audio to guide lip synchronization; when disabled, synchronize to the audio without a transcript. Test both settings on your own track, especially with fast delivery or music under the voice.enable_safety_checker, on by default. Content safety checks run unless you switch them off.seed, random if you omit it. The response hands back the seed it used. Log it, because it is the only handle you have on a take you liked.resolution, one of 480P, 768P, 1080P or 2K, with 768P as the default.
What one talking-head clip costs at each resolution
The pricing block on the model page is two sentences long, and the first one is the whole model: billing is calculated per second of generated video.
Per second of output, at four tiers
As of September 18, 2026, the rates are $0.05 per second at 480p, $0.08 at 768p, $0.16 at 1080p and $0.32 at 2K. fal's own worked example: a 5-second video at 768p costs $0.40. Across the lengths the audio rules allow, that gives:
| Clip length | 480p | 768p | 1080p | 2K |
|---|---|---|---|---|
| 5 s, the floor | $0.25 | $0.40 | $0.80 | $1.60 |
| 10 s | $0.50 | $0.80 | $1.60 | $3.20 |
| 15 s, the ceiling | $0.75 | $1.20 | $2.40 | $4.80 |
Every step above 768p doubles the rate. For a face watched on a phone, 768p or 1080p is where most of the money belongs. 2K earns its doubling over 1080p on a large display, or when the still carries a label whose fine print has to stay readable.
List price on day one, while its siblings run at half
This is the dated detail a budget should not miss. Five of the eight H3 Max endpoints are on sale. Text-to-video, image-to-video, Camera Controls and both Turbo entries carry promotional launch rates, 50% off for a limited time, and fal dates the end: the discount ends September 30. After it, image-to-video goes to $0.05, $0.08 and $0.16 per second at 480p, 768p and 1080p.
Those are the same three numbers Lip Sync charges today. Its pricing text is the four rates and the worked example, with no promotional note beside them. So Lip Sync launched at image-to-video's list rate, not its sale rate. The image-to-video entry prices three tiers, from 480p to 1080p; Lip Sync prices a fourth, 2K at $0.32 per second. Two practical consequences:
- A budget built from a sibling's promo rate is off by half for lip sync.
- The premium for lip sync disappears from 480p to 1080p once the siblings' discount ends on September 30 and image-to-video returns to the same three rates.
The floor is the audio, not a minimum charge
Per-second pricing on fal does not always mean per second. Director, the realtime sibling, bills each session at a minimum of 60 seconds runtime, and the billing increment under a per-second sticker is exactly where Mirage Avatar X's two price tags came apart.
Lip Sync's floor comes from the input rule instead. Five seconds of audio is the shortest clip, so the smallest bill is $0.25 at 480p and $0.40 at 768p. The response's duration field is the number to reconcile against your invoice, because it is the length fal actually generated.

Building a 30-second ad around a 15-second ceiling
Fifteen seconds holds a hook and a claim. It seldom holds a demonstration, an offer and a call to action as well, which is what a 30-second spot is for.
A 30-second spot is two calls
A 30-second talking head is two generations from the same still, each carrying half the voiceover: $2.40 at 768p, $4.80 at 1080p. A 60-second explainer is four. Each call stands alone. The six input fields are the image, the audio, the resolution, the transcription switch, the seed and the safety switch, and none of them refers to a previous clip.
So nothing carries the head position, the blink or the lighting from the end of one clip into the start of the next. Expect a visible seam at every join, and decide where it falls before you split the audio. Blink, light and skin are also where a synthetic take gives itself away one clip at a time, which our checklist for making AI video look real works through field by field.
Cut the script where a cut can hide
- Split on a full stop, never mid-phrase. Leave a beat of room tone at each end so the mouth starts and finishes closed.
- Keep both halves at five seconds or more. The audio floor applies per call, so a 17-second script splits 9 and 8, not 15 and 2.
- Cover the join with a product shot. The lip-sync endpoint gives you a face. The pack shot comes from somewhere else: a real photo, or a separate image-to-video clip.
- Or cut on purpose. Talking-head UGC is jump-cut by habit, and a hard cut on a new sentence reads as editing rather than as an error.
Then the assembly starts: captions, a hook card, an export per placement. A lip-synced clip is still raw material, and a clip is not an ad until somebody builds one out of it.
The call to action is the part that gets clipped
Clipping keeps the first 15 seconds, and scripts put the ask at the end. So an overlong voiceover does not fail loudly. It succeeds with the call to action removed, which is the worst way for an ad to fail, because the result still looks finished.
The pre-flight check is one line. Measure the file before you send it:
ffprobe -v error -show_entries format=duration -of csv=p=0 voiceover.mp3
- Over 15.0: split it at a sentence break before upload.
- Under 5.0: extend the line or merge it with the next one.
- How you know it worked: the
durationin the response equals the length you measured.
Worked example: ten hooks and one hero cut for a serum
Ad testing is bought by the variant, not by the clip, so here is the arithmetic on a typical testing brief.
The brief
A skincare brand selling a vitamin C serum wants to test ten opening lines. The setup:
- One 9:16 still of the founder holding the bottle, so the product is in every frame.
- Ten 15-second voiceovers. The last ten seconds are identical and only the first five change.
- One 30-second hero version for the retargeting audience, built as two calls.
Ten is a sensible opening set for one concept, though the number you actually need depends more on budget and event volume than on taste.
The bill at three resolutions
| What you render | Calls | Seconds | 768p | 1080p | 2K |
|---|---|---|---|---|---|
| Ten 15 s hook variants | 10 | 150 | $12.00 | $24.00 | $48.00 |
| One 30 s hero cut | 2 | 30 | $2.40 | $4.80 | $9.60 |
| Total | 12 | 180 | $14.40 | $28.80 | $57.60 |
Now suppose three of the ten come back with sync you would not run and need a second take. At 1080p that adds $7.20, for $36.00 all in. For a tested set of talking heads, that is a small number. It is also not the whole bill.
What the arithmetic cannot price
- The voice. Ten reads from a person, or ten synthetic renders, paid for wherever you make them.
- The assembly. Joins, captions, hook cards and exports are editing time, not fal time.
- The sync quality. The endpoint went public on the morning this post was written, and this research found no independent evaluation yet of how it handles fast speech, hard consonants or non-English audio.
- The speed. fal's engineering post from late on September 17 opens by saying H3 Max generates a 5-second video in under 3 seconds, but it uses the image-to-video endpoint as its example call and went up before Lip Sync was public. Treat that as a number about the family, not a promise about this endpoint.

Lip-sync endpoint or talking actor: which fits the job
Novoads competes in this space: it makes talking-head video ads too, so read this section knowing that. Novoads does not run H3 Max Lip Sync, and there is no MiniMax or H3 model in its catalog. What follows compares two routes to the same shot.
Same photo, a different input
The closest thing in Novoads to H3's still-to-video step is a custom actor. Under Create Actor > Upload, the app's own line is "Transform a picture into a talking actor." The difference is what you hand over next. H3 wants an audio file. The talking actor wants a script of up to 1,500 characters, and it generates the voice from a catalog with voices in 31 languages.
The product-in-hand rule works the same way on both sides, for the same reason. A stock library actor in Novoads has empty hands, so to show the product held on camera you create a custom actor from a photo of someone holding it, or start from a Discover product template, which swaps your uploaded product image into a template creator's hand. On fal, the product is in the clip only if it is in your still.
What each one bills
| H3 Max Lip Sync on fal | Novoads talking actor | |
|---|---|---|
| Starts from | A still and an audio file | A photo or library actor, plus a script |
| Voice | You supply it | Generated from the script |
| Longest input | 15 s of audio per call | 1,500-character script |
| Output resolution | 480p to 2K | 720p |
| Billing unit | Per second, by resolution | Per started minute, per actor |
| One 15 s clip | $0.75 to $4.80 | 10 credits |
| One 60 s spot | Four calls, $3.00 to $19.20 | 10 credits, one render |
Novoads starts at $49/month (Starter, 50 credits per month), so 10 credits is a fifth of a Starter month, or $9.80 at that plan's rate. Read the table straight and the verdict splits by length. For a single 15-second clip, fal is cheaper at every resolution, and at 1080p and 2K it is sharper. For a full minute at 1080p, fal's $9.60 is about the same money, before the voice track and three joins. Comparing a per-second sticker with a per-credit one is its own skill, and how AI video credits translate into dollars walks through it.
Use H3 Max Lip Sync when, use a talking actor when
Use H3 Max Lip Sync when:
- You already own the voice: a founder recording, a licensed voice actor, a dub.
- The deliverable is 5 to 15 seconds and resolution matters.
- You are wiring your own pipeline against an API and want to pay by the second.
Use a talking actor when:
- You start from a script, not a recording.
- The spot runs past 15 seconds and you would rather not cut joins.
- You want to cast from a library of 100+ AI actors and export from the same place.
How Novoads solves the script-to-talking-head step
Novoads turns a typed script into a talking-head video ad. You pick one of 100+ AI actors or upload a photo as a custom actor, choose a voice, and the lip-synced video comes back in the same workspace where you add captions, billed at 10 credits per started minute.
The honest pitch is narrow. It takes the audio file, the 15-second split and the joins out of the talking-head job, at 720p rather than 2K. If that trade suits the ads you run, Novoads starts at $49/month on the Starter plan, with 50 credits every month and cancellation whenever you want.

The voice track is the ad; the face is the delivery
H3 Max Lip Sync is a clean piece of engineering with its limits printed on the page: a five-second floor, a fifteen-second ceiling, four resolutions, and a per-second rate with no discount waiting to expire. It animates exactly what you hand it, and it hands back exactly that soundtrack.
That moves the whole creative decision upstream, into the recording. The endpoint will not rescue a flat read, a buried hook or an offer that lands at second nineteen. Write the fifteen seconds first, measure them, then pay for the face. A lip-sync model makes any line look believable, which is exactly why the line has to be worth believing.
Frequently Asked Questions
What is H3 Max Lip Sync?
It is an image-to-video endpoint on fal, made public on September 18, 2026. You send one still image and one audio file, and it returns a video in which the face in the image speaks the audio. fal's own description reads: H3 Max Lip Sync generates a video from an image and supplied audio, synchronizing mouth movements to the soundtrack. It supports optional transcription guidance and output resolutions from 480p to 2K.
Is H3 Max Lip Sync a MiniMax release?
No. The endpoint sits under a minimax/ path and fal's catalog files it under model lab Minimax, but fal's own H3 Max page states that H3 Max is post-trained by fal on top of the open-weight base MiniMax H3 model, and that standard MiniMax H3 is a separate frontier model with its own endpoints. MiniMax published the open weights; fal post-trained them and built the lip-sync endpoint.
How much does H3 Max Lip Sync cost?
As published on September 18, 2026, billing is calculated per second of generated video: $0.05 at 480p, $0.08 at 768p, $0.16 at 1080p and $0.32 at 2K. Because the audio must run 5 to 15 seconds, one call costs between $0.25 (5 seconds at 480p) and $4.80 (15 seconds at 2K). Unlike five of its H3 Max siblings, it launched without a 50% promotional rate, so there is no discount date to watch.
How long can the audio be?
At least 5 seconds, and audio over 15 seconds is automatically clipped to its first 15 seconds. The output video matches the clipped audio length, so a longer voiceover loses its ending, which is usually where the call to action sits. A 30-second spot has to be built from two calls and joined in an edit.
What aspect ratios and resolutions does it support?
The input image's aspect ratio must be between 0.4 and 2.5, which covers 9:16, 4:5, 1:1 and 16:9. Output resolution is 480P, 768P, 1080P or 2K, with 768P as the default, and the endpoint uses the supported aspect ratio nearest your image, so crop the still to your placement before you upload it.
Can I use H3 Max Lip Sync inside Novoads?
No. Novoads does not run H3 Max Lip Sync or any H3 model. Its own route to the same shot is the talking actor: pick a library actor or upload a photo as a custom actor, type a script, and the voice is generated from it. A talking-actor render bills 10 credits per started minute at 720p, and Novoads starts at $49/month (Starter, 50 credits per month).
Key Takeaways
- fal made H3 Max Lip Sync public on September 18, 2026, stamped at 05:26 UTC in its own catalog. One still image and one audio file go in, and a lip-synced video carrying that same soundtrack comes out. It is the eighth H3 Max endpoint fal has published since August 26.
- H3 Max is fal's post-train of MiniMax's open-weight H3. The endpoint path starts with minimax/ and the catalog files it under model lab Minimax, but fal built it; MiniMax published the base weights.
- Billing is per second of generated video, as of September 18: $0.05 at 480p, $0.08 at 768p, $0.16 at 1080p and $0.32 at 2K. It launched at those rates with no promotional discount, while five sibling H3 Max endpoints run at 50% off until September 30.
- Audio must be at least 5 seconds and anything past 15 seconds is clipped to its first 15. One call therefore costs $0.25 to $4.80, a 30-second spot is two calls joined in an edit, and an overlong voiceover silently loses its ending, which is usually the call to action.
- Novoads does not run H3. Its talking actor starts from a typed script instead of an audio file, renders at 720p and bills 10 credits per started minute. fal is cheaper and sharper for a single 15-second clip; the talking actor covers a full minute in one render with the voice generated from the script.


