Skip to main content

MiniMax H3 Max Styles: VHS, 70s Toon and Pixel-Art Ads at $0.08 a Second

fal now lists five one-call H3 Max style endpoints, VHS, Retro Toon 70s, Hand Drawn, Low Poly and 16-bit Pixel, each making 5 to 15 seconds of 768p video with sound from a prompt or an optional first frame at $0.08 a second. Here is which look fits which ad, what a test costs this week, and how to brief one.

Mauricio Valdivia

Mauricio Valdivia

·12 min

A creative desk with a VHS cassette, a small CRT television glowing with static, pencil animation sketches and a phone filming a cold brew can

The style is picked before you write the prompt

A cold brew brand wants its next hook to look like a Saturday-morning cartoon from 1974. Until recently that meant a paragraph of style direction, a few rerolls, and a look that could shift from take to take. Now the style has its own endpoint.

fal now lists five one-call H3 Max style endpoints: VHS, Retro Toon 70s, Hand Drawn, Low Poly and 16-bit Pixel. Each makes 768p video with audio, 5 to 15 seconds long, from a text prompt or an optional first-frame image. Each bills the same way: $0.08 a second with the sound included, so a 5-second clip costs $0.40 and a 15-second clip costs $1.20.

The look is no longer the hard part. The idea still is.

Below: what fal lists and what is actually new, what a preset takes out of your brief, what a stylized test costs this week, which look fits which ad, how to brief one, and how to build the same kind of spot with the models Novoads runs, which do not include H3.

What fal lists under /styles/

The five model pages carry no date. What they do carry is a contract, and it is the same contract five times.

Five endpoints, one contract

All five live under one path, minimax/h3-max/styles/, and each page opens with the same shape of sentence: it generates 768p video with audio from text prompts or an optional first-frame image. The input schemas match field for field, except for one extra switch on VHS. Read together, the shared terms are:

  • Resolution: 768p. The pricing line prices one tier, and the output field describes a 768P video with audio.
  • Length: 5 to 15 seconds, set as a whole number. The default is 5.
  • Inputs: a prompt, which is required, and a first-frame image, which is optional.
  • Frame: six aspect ratios for text-only calls, 21:9, 16:9, 4:3, 1:1, 3:4 and 9:16, with 16:9 as the default. With a first frame, the image's ratio sets the canvas and the aspect setting is ignored.
  • Sound: generated with the picture and included in the per-second price.
  • Output: the video and the seed that made it.

What each style is for, in fal's words

Each page describes its style in a few words and recommends a kind of first-frame image. Those recommendations are the most practical line on the page, because they tell you what the model expects to start from.

Endpointfal's descriptionFirst frame fal suggests
VHSVHS-style, adjustable tape damageA 4:3 image
Retro Toon 70s1970s hand-painted animationPainted-style artwork
Hand DrawnHand-drawn animationPencil artwork
Low PolyRetro low-poly 3DLow-poly or retro 3D artwork
16-bit Pixel16-bit pixel artPixel artwork

VHS gets the only extra control, damage_level, with three settings. fal describes them this way: light adds subtle analog noise, medium adds visible damage, and heavy adds strong tracking errors and distortion. Medium is the default, so a call that never mentions damage still gets visibly worn tape.

What is new here, and what is not

fal's catalog is where the dates live. It carries two stamps for each of the five: an entry created on the evening of September 17, 2026 (UTC), and a public stamp of September 21. It files all five under model lab Minimax and family H3, but the post-train underneath is fal's. In fal's words on its H3 Max page, H3 Max is post-trained by fal on top of the open-weight base MiniMax H3 model. Our explainer on what MiniMax H3 is covers the base model.

VHS is the one look with a history. A style LoRA called vh5tape, described on its Hugging Face card as a worn-VHS look for MiniMax H3 with three trained damage levels, has been on Hugging Face since August 31. Its card says it runs on fal's H3 LoRA endpoints, which fal's catalog shows as public since August 10. That route asks for a weights file, a scale setting and a trigger word at the start of the prompt.

So the new part is packaging. A style endpoint takes a prompt and returns the look, and tape wear becomes a three-way switch instead of a phrase you have to remember. For an ad team, that is the difference between a technique and a tool. The five styles also join a family of H3 Max endpoints fal has been publishing since late August, from the realtime H3 Max Director to H3 Max Lip Sync.

Several UGC creators filming product variations to camera
Novoads · UGC video ads with AI, ready in minutes.
Try now

What a preset takes out of your brief

A style endpoint does to the look what a template does to a layout. It removes the decision, and with it the drift. It also removes some controls, and one of them matters more than the rest for ads.

The prompt field changes job

Every style's prompt field carries the same instruction, with only the style name changing: "Describe the scene, action, camera and sound." It then says the style is applied automatically. On VHS the wording is that VHS style and tape damage are applied automatically.

Our reading of that sentence is simple. Words spent on the look are words taken from the shot. "Retro 1970s cartoon, hand-painted, grainy palette" tells the endpoint nothing it was not already going to do. "A barista slides a can down a diner counter, wide shot, slow push-in, jukebox music" tells it everything it cannot guess.

Four of fal's five page samples are written that way. They read like shot notes: a subject, an action, a camera instruction such as "Medium locked shot" or "steady tracking camera", and a line of sound such as "Bicycle chain clicks and gentle market ambience." Four of the five samples on the pages carry a sound line, and two end with the words "no speech". The exception is VHS, whose page sample opens on the look itself: "Badly damaged VHS tape with heavy tracking errors and distortion. A 1980s horror movie on a rented tape".

What you give up next to plain H3 Max

The comparison is exact, because both schemas are complete lists. Plain H3 Max text-to-video takes nine input fields. A style endpoint takes five, or six on VHS.

Plain H3 Max text-to-video gives you:

  • A resolution choice of 480P, 768P or 1080P. fal describes the 1080P option as latent refinement from a native 768P source.
  • Your own soundtrack, through a field for an audio clip at least 2 seconds long and up to 15 MB, pinned to the generated soundtrack.
  • A prompt-expansion setting: disabled, balanced, or quality, which spends up to about 30 seconds on a richer prompt.
  • A safety-checker switch, on by default.

A style endpoint gives you instead:

  • A trained look you never have to describe.
  • An optional first frame in the same call. Plain text-to-video has no image field; that job belongs to a separate image-to-video endpoint.
  • The wear switch, on VHS only.

The missing audio field is the one to plan around. On a style endpoint the sound is always generated. That is fine for ambience and effects, and a problem for a brand voiceover, a licensed track or a jingle people already know, which have to go on in the edit.

The resolution gap is smaller than it looks. If 1080P on the plain endpoint is a refinement of a native 768P render, then 768p is H3 Max's top native tier, and the styles simply skip the refinement step. For a VHS or pixel look, softness is part of the style anyway.

The frame defaults to landscape

Forget one parameter and you get a 16:9 clip. For Stories, Reels and TikTok, set aspect_ratio to 9:16 on every text-only call. The list has no 4:5, so a feed ad either takes 3:4 or 1:1, or starts from a 4:5 first frame, since fal says the image's ratio decides the canvas. Check the returned file either way, and measure it against the spec each platform publishes.

VHS adds a trap of its own. Its first-frame field recommends a 4:3 image for the classic VHS frame. That is the period-correct shape, not necessarily the one your placement is bought in. Pick the placement first and let the tape wear carry the period.

What a stylized clip costs, this week and after September 30

The style pages price one tier, and the rate is identical on all five: billing is calculated per second of video generated, and video at 768p costs $0.08 per second with audio included.

One rate, sound included

At that rate the three lengths an ad team actually uses come to:

  • 5 seconds: $0.40, which is also what a call with no duration set costs, since 5 is the default.
  • 10 seconds: $0.80.
  • 15 seconds: $1.20, the ceiling for one call.

A 30-second spot is two calls and a join. Nothing in the input schema refers to a previous clip, so decide where the cut falls before you write the prompts. Our suggestion for the seam: take the last frame of the first clip and pass it as the first frame of the second. It is already artwork in the right style, which is exactly what fal asks a first frame to be.

Twice the plain endpoint until September 30

This is the dated detail a budget should not miss. As of September 23, plain H3 Max text-to-video charges $0.04 per second at 768p. fal marks that as a promotional launch rate, 50% off for a limited time, and the discount ends September 30, after which 768p goes to $0.08. H3 Max Turbo text-to-video sits lower again. The style pages carry no promotional note.

Endpoint at 768pRate on Sept 23Rate fal lists after Sept 30
Any H3 Max style$0.08/sNo change listed
H3 Max text-to-video$0.04/s$0.08/s
H3 Max Turbo text-to-video$0.02/s$0.04/s

So until September 30, a stylized second costs double a plain one. After the discount ends, the two cost the same at 768p, and only Turbo stays cheaper.

What the premium buys

Until September 30, the extra $0.04 a second buys a trained look instead of a described one. On a look test that premium is a rounding error: at the September 23 rates, 25 seconds of drafts cost $1.00 more than they would on the plain endpoint. On a 200-clip library the question changes. Until September 30 the plain endpoint renders the same seconds for half, if words alone can hold your style. After that date, price stops being a reason to pick either one at 768p.

Which look fits which ad

Nothing on fal's pages says which style sells what. What follows is our reading, built from what each look borrows and what it costs the product on screen.

Tape and toon: looks that borrow a memory

VHS works when the ad wants to feel found rather than made. Light wear suits a throwback testimonial or a "remember this?" hook for a brand with a long history. Heavy wear suits horror, a Halloween drop, or a parody of a late-night infomercial, and it will chew through a label, because strong tracking errors and distortion are the point.

Retro Toon 70s suits a mascot, an origin story or a hook that needs scale nobody could afford to shoot. fal's own sample is a warrior on a ridge while a dragon glides past below. Hand Drawn suits a gag: fal's sample has a weary office worker whose neck stretches upward in surprise when a bird lands on his head, a squash-and-stretch joke no live-action shoot delivers cheaply.

Polygons and pixels: looks that borrow a game

Low Poly suits games, apps and tech launches, and any hook built on motion. fal's sample is a woman in a red coat running across a rainy rooftop at night, with the camera tracking alongside her. 16-bit Pixel suits gaming, retro merchandise and apps with a sense of levels or progress, and it is the easiest of the five to turn into a joke about the product "powering up".

LookAd job it suitsWatch for
VHS, lightThrowback and heritage hooksWear softens fine print
VHS, heavyHorror, Halloween, parodyTracking errors eat labels
Retro Toon 70sMascots, big-scale hooksNeeds a painted first frame
Hand DrawnGags and quick explainersFlattens product texture
Low PolyGames, apps, tech launchesFacets break round packaging
16-bit PixelGaming, retro merch, appsSmall logos vanish

When not to use a preset

A preset is a hook device, not a product shot. Skip it when:

  • The texture sells the product. Skincare, food and fabric live on close-up detail that every one of these styles softens or simplifies on purpose.
  • The label has to be read. Pack shots and legal lines belong in a clean frame.
  • The ad runs on a real-looking face. A testimonial borrows trust from looking unproduced, and a cartoon spends that trust.

Where a preset earns its place is variety. Meta says its ads system groups together ads that share key visual and thematic attributes, and that ads truly distinct in their visuals, messaging and formats expand the range of options it can use to reach different audiences. Our breakdown of creative diversity in Meta ads walks through that guidance.

Our reading: a pixel cut and a live-action cut of one script share a message but not a look, which makes a style swap one of the cheaper ways to change the visuals. Whether Meta's system counts it as distinct is Meta's call, not ours. Test it the way you would test ten hook variations: one change at a time.

A UGC creator holding a product up to the camera
Novoads · UGC video ads with AI, ready in minutes.
Try now

How to brief a style endpoint

The brief is short because the endpoint already owns the look. What is left is the shot, the product and the sound.

Write the shot, not the style

Use the order the prompt field asks for: scene, action, camera, sound.

[Subject] [does one clear action] in [place]. [Shot size], [camera move].
[Sound line], [no speech].

For the cold brew brand on Retro Toon 70s, that becomes: "A cartoon barista slides a can of cold brew down a long diner counter. It spins and stops upright in front of a sleepy customer, who jolts awake. Wide shot, slow push-in. Jukebox music and a metallic clink, no speech." It is one gag with one payoff, which is about what a five-second draft can hold. Direction words that land and the ones that break a shot are in our guide to camera movement prompts for ads.

Put the product in the first frame, drawn in the style

fal's first-frame advice is the same on every page: start from artwork that already matches the look. Our reading is that a studio packshot fights the preset. The endpoint is built to return pencil lines or pixel clusters, and your photo asks it to hold something else.

So make the first frame first. Generate a still of your product in the target style with any image model, keep the label large and simple, crop it to your placement, and pass it as image_url. Its ratio sets the canvas, which solves the aspect question in the same move. This is the same two-step shape as our guide to 3D animated product ads: the still decides the look, the render decides the motion.

Plan the sound as generated

There is no audio field, so write the sound you want into the prompt and treat the result as a guide track. Ambience, effects and music beds are what fal's samples ask for. None of the samples on the five pages includes a spoken line, so if a character has to say your tagline, test that on one clip before you script ten around it.

How to know it worked

  • The frame matches the placement: 9:16 was set, or the first frame was already cropped.
  • The product reads at phone size on the opening and closing frames.
  • The sound line you wrote is audible, with no stray speech.
  • The seed is logged. The response returns it, and it is your only handle on a take you want again.

Worked example: one cold brew hook in five looks

Ad testing is bought by the variant, so here is the arithmetic on a typical look test.

The brief

A cold brew brand has one hook concept and no fixed look. It wants to find the one or two styles worth a full spot, then cut three opening lines in each. Drafts run at the default 5 seconds; finals run the full 15.

  • Round 1: the same hook in all five styles, 5 seconds each. The VHS draft runs at the default medium wear.
  • Round 2: the VHS draft again at light and heavy, so all three wear levels sit side by side.
  • Round 3: the two winning looks, three hooks each, at 15 seconds.
  • Rerolls: one more take on three of those finals.

The bill

RoundRendersSecondsAt $0.08/s
Five styles, 5 s drafts525$2.00
VHS wear, light and heavy210$0.80
Two looks x three hooks, 15 s690$7.20
Rerolls on three finals345$3.60
Total16170$13.60

The same 170 seconds on plain H3 Max text-to-video would cost $6.80 until September 30, with the style written into the prompt by hand, and $13.60 after it. For a look test, the preset is cheap insurance against a hand-written style that drifts between takes.

What the bill leaves out

  • The styled first frames, made and paid for wherever you make images.
  • The edit: captions, an end card with a clean pack shot, and any voiceover or licensed music laid over the generated sound.
  • The 30-second cut, which is two 15-second calls joined in the edit.
  • Your own quality bar. fal's pages show fal's samples, not your product. Judge the five looks on a can with your label on it.

How Novoads solves the stylized product ad

Novoads is a video-ad tool, so weigh this section knowing that. It turns a product photo into ad images and video with the image and video models it runs, in one workspace. It does not run H3 or any of these style endpoints, and it has no VHS switch.

The route is the two-step one from the briefing section. First, a styled still of the product from one of the image models Novoads runs: GPT Image 2, GPT Image 2.5 Sunburst and Flare, Nano Banana Pro, Seedream 5 Lite and Pro, and Reve 2.1. Then an image-to-video render that starts from that still, such as Seedance 2.0, the route our 3D product-ad guide uses. The video roster also covers Seedance 2.5, Seedance 2.0 Mini, Kling v3 Pro, Kling Motion Control, Google Veo 3.1 and Omni Flash.

The honest difference is where the look lives. On fal, the endpoint holds it. In Novoads, the still holds it, and your words keep it. If that trade suits the ads you run, you can try the route in Novoads, and every plan is published on the pricing page.

A UGC creator filming a skincare product review on a phone
Novoads · UGC video ads with AI, ready in minutes.
Try now

Nostalgia is a hook, not a product shot

Five styles for one per-second rate is a clean piece of packaging. On VHS, it turns what the LoRA route did with a weights file, a trigger word and a scale into one call and one switch. It also moves every remaining decision onto the parts the preset cannot make: the scene, the one action, the sound, and whether the product survives the style.

That is the real brief now. A preset makes the look cheap. It does not make the idea good, and it cannot make a label legible through heavy tape damage. Use the style to stop the scroll, then let the product show up clean.

Frequently Asked Questions

What are the H3 Max style endpoints?

They are five text-to-video endpoints on fal, each with one visual style built in: VHS, Retro Toon 70s, Hand Drawn, Low Poly and 16-bit Pixel. You send a prompt describing the scene, action, camera and sound, and optionally a first-frame image, and the endpoint returns a 768p video with audio in that style, 5 to 15 seconds long. H3 Max itself is fal's post-train of MiniMax's open-weight H3 model.

How much do the H3 Max styles cost?

As listed on September 23, 2026, billing is per second of video generated: $0.08 per second at 768p with audio included, so a 5-second clip costs $0.40 and a 15-second clip costs $1.20. The default duration is 5 seconds. Plain H3 Max text-to-video costs $0.04 per second at 768p until its 50% launch discount ends on September 30, when fal lists it at $0.08.

Can I render the styles at 1080p or in 9:16?

Not at 1080p: the style endpoints output 768p and have no resolution field. Vertical works. For a text-only call you choose from 21:9, 16:9, 4:3, 1:1, 3:4 and 9:16, and the default is 16:9, so set 9:16 yourself for Stories, Reels and TikTok. When you pass a first-frame image, fal says the image's aspect ratio decides the output canvas instead.

Can I use my product photo or my own soundtrack?

A first-frame image is optional, and fal advises artwork that already matches the style: pixel artwork for 16-bit Pixel, pencil artwork for Hand Drawn, painted-style artwork for Retro Toon 70s. There is no field to pin your own audio, unlike plain H3 Max text-to-video, which takes an audio clip for its soundtrack. On a style endpoint the sound is generated, so a brand voiceover or licensed track goes on in the edit.

Is the VHS look new?

The endpoint is new; the look is not. A worn-VHS style LoRA for H3 with three trained damage levels has been on Hugging Face since August 31, 2026, and its card says it runs on fal's H3 LoRA endpoints. The VHS style endpoint packages that kind of look as one call with a damage_level switch of light, medium or heavy.

Can I use these styles inside Novoads?

No. Novoads does not run H3 or any H3 style endpoint. To make a stylized product ad in Novoads, generate a styled still of your product with one of its image models, then animate that still with an image-to-video model such as Seedance 2.0. Plans are published on novoads.ai/pricing.

Key Takeaways

  • fal now lists five H3 Max style endpoints under minimax/h3-max/styles/: VHS, Retro Toon 70s, Hand Drawn, Low Poly and 16-bit Pixel. Each makes 768p video with audio, 5 to 15 seconds long, from a text prompt or an optional first-frame image. fal's catalog stamps all five as created on September 17 and public on September 21, 2026.
  • Every style bills $0.08 per second with audio included: $0.40 for 5 seconds, $1.20 for 15. As of September 23 that is twice the $0.04 per second plain H3 Max text-to-video charges at 768p, because the plain endpoint runs at 50% off until September 30 and the styles carry no promotional rate.
  • A preset removes the style from your brief. Each prompt field asks you to describe the scene, action, camera and sound, and applies the look automatically. The trade is fewer controls: no resolution choice beyond 768p, no field to pin your own soundtrack, and a 16:9 default you have to change for vertical placements.
  • VHS is the one look with a history. A worn-VHS style LoRA for H3 with three trained wear levels has been on Hugging Face since August 31. What is new is the packaging: one call, the look applied for you, and tape wear reduced to light, medium or heavy.
  • Novoads does not run H3 or any style endpoint. The same kind of stylized product ad is built in two steps with the models it does run: a styled still of the product from an image model, then an image-to-video render that starts from it.
Mauricio Valdivia

Mauricio Valdivia

Founder of Novoads

Mauricio is the founder of Novoads, where he works to democratize video advertising with AI for brands in Latin America.

Related Articles

What is MiniMax H3 (Hailuo 3.0)? 2K video with native audio, explained for ad makers

What is MiniMax H3 (Hailuo 3.0)? 2K video with native audio, explained for ad makers

MiniMax launched H3 on July 31, 2026: one model that generates up to 15 seconds of 2K video with native stereo sound. Here is what those specs actually change for anyone making short-form video ads, and what they do not.

NewsMiniMaxAI video models
H3 Max Lip Sync on fal: a 15-second talking-head ad costs $0.75 to $4.80

H3 Max Lip Sync on fal: a 15-second talking-head ad costs $0.75 to $4.80

fal made H3 Max Lip Sync, its post-train of MiniMax's open-weight H3, public on September 18, 2026: it turns one still image and a 5 to 15 second voice track into a lip-synced clip billed at $0.05 to $0.32 per second, which prices a talking-head ad and caps every call at 15 seconds.

NewsAI video modelsLip sync
H3 Max Director streams AI video you can redirect mid-take, capped at two minutes

H3 Max Director streams AI video you can redirect mid-take, capped at two minutes

fal made H3 Max Director publicly callable on September 3, 2026. It generates one continuous video stream you steer with live prompts, at $0.02 per second of video until September 14 and a 60-second minimum per session.

NewsAI video modelsRealtime video
3D Animated Product Ads Without a Studio: One Still Image, One 15-Second Render

3D Animated Product Ads Without a Studio: One Still Image, One 15-Second Render

A stylized 3D animated product ad does not need an animation pipeline. Two generated stills decide the look and the staging, then one image-to-video call renders 15 seconds with the sound already in the file. Here is the route, the prompt rules that keep it from breaking, and what a finished spot costs.

Guides3D animated adsProduct ads

Ready to create video ads with AI?

Generate professional video ads in minutes, not weeks.

Start for $49/month