Skip to main content

What Is Flux 3? Black Forest Labs' Multimodal Model, Explained for Ad Makers

Black Forest Labs announced FLUX 3 on July 23, 2026: one model that generates images and up to 20-second video with native audio, and extends to action prediction. It is gated early access, not a launch. Here is what it changes for ad makers, and when you can actually use it.

Mauricio Valdivia

Mauricio Valdivia

·11 min

A production monitor showing a paused vertical product ad with an audio waveform, beside a small desktop robotic arm holding a cosmetics bottle

Video, Audio, and Action, From One Model

On July 23, 2026, Black Forest Labs announced FLUX 3, and the pitch is much bigger than a new image model. FLUX 3 is one foundation model that jointly learns from images, videos, and audio, generates clips up to 20 seconds long with the soundtrack already inside, and extends to predicting physical actions for robots. Not three models sharing a brand. One architecture.

The catch sits in one verb. BFL announced FLUX 3; it did not launch it. FLUX 3 Video and FLUX 3 Action are early access behind an application form, FLUX 3 Image opens in the coming weeks, general availability comes some unspecified time after that, and no price for any of it has been published.

For people who make ads, that mix of a genuinely new model shape and firmly gated access is exactly why the story deserves ten minutes. This piece walks through what BFL actually said, what you can and cannot do with FLUX 3 today, and what a unified video-and-audio model will change for ad creative once it stops being a waitlist.

What Black Forest Labs Actually Announced

Strip the world-model philosophy out of the announcement and three concrete things remain: a model, a rollout ladder, and a set of self-reported results. Each deserves its own reading.

One architecture, three modalities

In BFL's own words, "FLUX 3 is our new multimodal foundation model. It jointly learns from images, videos, and audio within a unified architecture." The company's argument for joint training is physical rather than mystical: the sound has to match the impact, the motion has to obey the mass, so a model that learns all three signals together learns more about the world than three specialists learn apart.

That framing is what separates this from a routine version bump. FLUX went from an image family, FLUX.1 and FLUX.2, to a single system that generates video with sound, synthesizes and edits images, and can be finetuned to drive robot arms. The launch release compresses it into one line: "jointly trained across image, video, audio, and action prediction modalities within a unified architecture."

Announced is not launched

Every FLUX 3 capability ships through a gate. The company's launch plan says it plainly: "Over the next few weeks and months, we will make the following capabilities available, each after an early access phase." Here is the ladder as it stood on announcement day:

PieceWhat BFL says it isStatus on July 23, 2026
FLUX 3 Videovideo with native audioearly access, by application
FLUX 3 Actionaction prediction for roboticsearly access, partner route
FLUX 3 Imageimage synthesis and editingopens in coming weeks
FLUX 3 Devopen-weight multimodal backbonelater this year
General availabilityfull rollout plus benchmarksto follow, no date

The last two rows carry no dates beyond "later this year." Plan around that, not around the headlines. FLUX 3 is not even alone at this stage: ByteDance's announced Seedance 2.5 sits in the same gated, waitlist-shaped position, down to the promise of longer single-pass clips nobody outside can run yet.

Day-one coverage is one press release wearing many mastheads

One detail worth knowing before you read about FLUX 3 anywhere else: the launch-day articles circulating on news sites are syndications of BFL's own GlobeNewswire press release, reprinted word for word. That does not make the facts wrong, but it means every day-one claim, including the flattering ones, traces back to the company. Independent testing does not exist yet, and BFL itself says full benchmark results and methodology arrive "alongside broader availability."

A UGC creator filming a skincare product review on a phone
Novoads · UGC video ads with AI, ready in minutes.
Try now

What FLUX 3 Video Promises for Ad Creative

The video piece is the part ad makers care about, so here is the spec sheet as BFL states it, translated into ad terms. Everything in this section is the vendor's description of a model still in development; nobody outside the early access program can verify it yet.

Twenty seconds, sound included

BFL says "FLUX 3 can create highly diverse videos with audio up to 20 seconds in length in a single generation," and that all outputs come with native audio generation. No silent clip you score afterward: the model writes the fizz, the thud, and the dialogue in the same pass as the pixels, in multiple languages.

Twenty seconds is also a meaningful number for advertising specifically. Most performance clips run 15 to 30 seconds, so a single generation covers a full ad body rather than a single beat. Compare that with the current crop of engines that top out around 10 or 12 seconds per generation and force you to stitch. The nearest active fight on both fronts is the native-audio and long-take race between Kling 3.0 and Seedance 2.5, and even there only one side ships today.

Five doors into a clip

FLUX 3 Video takes more than a text prompt. BFL lists five distinct input modes, and each one maps to a different job in an ad workflow:

  • Text-to-video: the blank-page mode, for B-roll and scene-setting shots.
  • Image-to-video: animate a starting frame, or steer with reference images. This is the product-photo door, the input most ad tests actually start from.
  • Video-to-video: carry "central elements of a source video - for instance the same character - into a new scene or context." Remix a winning ad without reshooting it.
  • Keyframe-to-video: define the first and last moment, let the model fill the transition. Storyboards become briefs.
  • Audio-video continuation: extend an existing clip and its sound instead of starting over.

For an ad team, the reference doors are the interesting ones. A product still or a source clip going in means the output stays anchored to your actual asset instead of the model's imagination of it. That is the difference between a usable product ad and a pretty hallucination.

Characters that survive the cut

The announcement also promises multi-shot construction: "agentic chaining of individual clips into longer, multi-shot sequences," with visual references keeping characters consistent across scenes in sequences that BFL says can last several minutes.

Consistency across cuts is the oldest failure mode in AI video ads. The same spokesperson has to appear in shot one and shot four, holding the same bottle, or the viewer's eye flags the whole thing as synthetic. If character carry-over holds up outside a demo reel, it is the single most ad-relevant line in the entire announcement.

The Numbers BFL Reports, and How to Read Them

The announcement carries percentages, and percentages travel fast while their caveats stay home. Pin them down before they show up unsourced in your feed.

A company grading its own homework

BFL reports that across early evaluations "FLUX 3 was preferred over Grok Imagine Video in up to 69% of comparisons, Kling v3 Pro in 60%, Happy Horse v1 in 59%, Happy Horse 1.1 in 57%, Seedance 2.0 and Gemini Omni Flash in 52%."

Read the spread before the headline. Against the strongest current engines, the preference is 52%: a coin flip with a thumb on the scale. And these are the vendor's own numbers, generated on 10-second 720p clips, explicitly labeled preliminary, with the methodology unpublished until broader availability. Preliminary, internal, and self-scored is not a benchmark. It is a claim, and it should be quoted as one. For a comparison of two engines on that list you can actually act on, the Seedance vs Kling head-to-head for UGC ads is built on spec sheets at identical pricing, not preference polls.

What BFL says the model is best at

The self-described strengths are at least usefully specific: capturing human facial expressions, associating sounds with physical events, multilingual capability, and "strong typography generation and animated designs."

Notice how precisely that list maps onto ad craft. Faces carry trust. Sound sells physicality. Languages carry markets. Typography carries the offer. Whether or not the percentages survive independent testing, BFL is aiming this model at commercial creative work, and the launch release says so directly by naming e-commerce and "the maintenance of product and material consistency across motion" among its target applications.

Self-Flow is March research in a July coat

The training approach underneath FLUX 3, Self-Flow, is not part of this news. BFL published Self-Flow as a research preview in early March 2026, more than four months before FLUX 3, as independent coverage from GIGAZINE recorded at the time. What is new in July is scale: BFL says it "significantly scaled up compute and data resources" on that method to train FLUX 3. If a writeup presents Self-Flow itself as a launch-day breakthrough, it is recycling the spring as summer news.

A UGC creator filming a product review without a film crew
Novoads · UGC video ads with AI, ready in minutes.
Try now

What You Cannot Do With FLUX 3 Today

This is the section that saves you a wasted afternoon. The gap between the announcement and the product is wide, and it is worth stating each plank of it exactly.

The application form is the only door

FLUX 3 Video and FLUX 3 Action are available by early access application; FLUX 3 Image is not available at all yet. There is no public API key to create, no playground to open, no way to run a test brief through it this week unless BFL approves your application. "Request early access" is the only button the announcement offers.

No price exists

BFL's pricing page is pay-as-you-go, and its calculator covers FLUX.2 models only. There is no FLUX 3 SKU, no per-second rate, no tier, nothing to build a cost model on. For a media buyer this matters more than any capability claim: you cannot compute a cost per usable ad for a model with no price, and any article quoting one today is guessing.

Not on the platforms ad tools rent from

Most ad tools do not call model vendors directly; they rent engines through aggregators. As of July 23, 2026, fal's public model explorer lists FLUX.1 and FLUX.2 family endpoints and no FLUX 3 endpoint at all. Until that changes, no ad platform built on rented engines can offer FLUX 3, whatever its marketing says. Never build an ad pipeline on a waitlist. Availability risk runs in the other direction too: Sora's slide out of general availability showed how fast a model ad teams rely on can become a moving target.

Three signals that the gate is actually opening

Rather than refreshing the announcement page, watch for the three events that would turn FLUX 3 from a story into a tool:

  • A price appears. The day BFL's pricing calculator gains a FLUX 3 entry, you can finally compute a cost per usable ad and compare it against the engines you run now.
  • An aggregator lists an endpoint. A FLUX 3 model on fal or a comparable host means ad platforms can start wiring it in, and real per-clip latencies and failure rates become public knowledge.
  • The benchmark methodology ships. BFL has promised full results and methodology alongside broader availability. Once independent testers can rerun the comparisons, the preference percentages become checkable.

Until at least one of those happens, the correct amount of FLUX 3 in your production pipeline is zero.

What a Unified Model Changes for Ad Makers

Assume the rollout lands the way BFL plans it. Why should someone who ships ads for a living care about the architecture at all? Three reasons, each concrete.

One pass replaces the stitch pipeline

Today's common workflow generates silent video with one model, voices it with a TTS engine, and layers music or foley in an editor. Every seam is a chance for the sound to disagree with the picture. It is the same assembly tax you pay when running a raw video model like MiniMax's Hailuo for ads: the engine makes the pictures, and everything around them is your job. A model that generates video and audio jointly removes those seams by construction: the pour sounds like a pour because the same network made both.

Run the arithmetic on a real testing week to see why that compounds. The seam equation: handoffs per week = variants x pipeline stages. Ten ad variants through a three-stage stitch pipeline is 30 handoffs, 30 places where a voice drifts out of sync or a sound effect lands a frame late, each needing a human check. The same ten variants through a unified generation is 10 handoffs, and the audio cannot disagree with its own frames. At the hundreds-of-variants scale where paid social testing actually operates, that difference is a head count.

Product consistency is the e-commerce tell

The single most commercial sentence in the launch release names "the maintenance of product and material consistency across motion." Anyone who has generated product video knows why that phrase exists: labels warp, bottle caps change color, fabric turns to liquid mid-rotation. A product that stays itself across 20 seconds of motion is the difference between AI b-roll you can run under a real offer and AI b-roll you quietly delete. BFL naming it as a design target says it is chasing exactly the ad and e-commerce budget.

The ceiling finally maps to the format

Twenty-second single generations, chaining into minutes, and multilingual dialogue line up suspiciously well with how paid social actually works: a hook test lives in 6 seconds, a full spot in 20, a localized campaign in five languages. Current engines force you to assemble that from short silent pieces. A unified model, at general availability, would collapse much of that assembly into prompting. That is the future worth planning for, without pretending it shipped this morning. Meanwhile the buying side raises the stakes: Google's new video campaign groups coordinate frequency across YouTube campaigns, and Google's own best-practice math multiplies the distinct creatives each advertiser needs.

A grid of real UGC creators filming product videos
Novoads · UGC video ads with AI, ready in minutes.
Try now

How Novoads Fits While FLUX 3 Sits Behind a Waitlist

A model-agnostic ad workflow treats every frontier release the same way: track it, price it when a price exists, and add it when it can actually be served. FLUX 3 clears none of those bars yet. What you can run today is already strong.

The engines you can run right now

Novoads generates UGC-style video ads on Seedance 2.0 and its half-price Seedance 2.0 Mini variant, Kling v3 Pro, Google Veo 3.1, and Sora 2 and Sora 2 Pro, plus a talking-actor engine for spokesperson clips with lip-sync and captions. You write or auto-generate a script, pick an AI actor whose age and accent match your buyer, and get a vertical 9:16 ad ready for TikTok, Reels, or Meta. A finished clip runs roughly $2 to $11 depending on the engine, which is what makes real variant testing affordable while the frontier models fight over preference percentages. And if the product sells without a presenter at all, faceless video ads, product b-roll, screens, and voiceover carried by captions, are among the fastest formats an ad team can ship while the frontier sorts itself out.

Try the workflow for $1

The trial is deliberately small and honest: $1 for 3 days of access, recurring, and it becomes the $49-a-month Inicial plan if you stay. The first charge grants enough credits for about one video, so you can judge the output quality on your own product before committing to anything. Cancel anytime. When FLUX 3 eventually reaches general availability with a real price, a model-agnostic workflow is also the fastest way to actually use it on ads, because swapping an engine underneath is the platform's job, not yours.

A Direction You Plan For, Not a Tool You Buy Today

FLUX 3 is the clearest statement yet of where visual AI is heading: one model that learns the picture, the sound, and the action together, from a lab with the track record to be taken seriously. The direction is real, and parts of the spec sheet, native audio in every clip, character consistency across scenes, product fidelity in motion, read like they were written by someone staring at an ad brief.

But a direction is not a deliverable. Today FLUX 3 is an application form, a pricing page without a price, and a set of self-graded percentages awaiting methodology. The right move for an ad team is the boring one: keep shipping variants on the engines that exist, and let the waitlist prove itself. When the gate opens, the advantage will not go to whoever read the most launch coverage. It will go to whoever already has a testing pipeline the new engine can drop into.

Frequently Asked Questions

What is FLUX 3?

FLUX 3 is Black Forest Labs' multimodal foundation model, announced on July 23, 2026. Instead of separate models for each medium, it jointly learns from images, videos, and audio within one architecture, generates video with native audio up to 20 seconds per clip, synthesizes and edits images, and can be extended to action prediction for robotics through FLUX-mimic.

Is FLUX 3 available to use right now?

Not generally. As of the announcement, FLUX 3 Video and FLUX 3 Action are in early access behind an application form, FLUX 3 Image opens its early access phase in the following weeks, and general availability comes after that. BFL also says faster and open-weight versions arrive later this year. There is no date for any of it.

How much does FLUX 3 cost?

Nobody outside BFL knows yet. As of July 23, 2026, BFL's pricing page and its pay-as-you-go calculator cover FLUX.2 models only; no FLUX 3 SKU or per-second rate has been published. Any FLUX 3 price you see quoted today is a guess.

Can FLUX 3 generate video with sound?

Yes, according to BFL. The company says FLUX 3 creates videos with audio up to 20 seconds in length in a single generation, and that all video outputs come with native audio, including multilingual dialogue. That claim is only testable by early access users for now.

Can I use FLUX 3 inside Novoads?

Not as of this writing. Novoads runs Seedance 2.0, Seedance 2.0 Mini, Kling v3 Pro, Google Veo 3.1, Sora 2, and Sora 2 Pro as video models, plus a talking-actor engine for spokesperson ads. FLUX 3 is gated early access at BFL and is not served by the aggregator platforms ad tools rent models from, so no ad platform offers it today. The workflow is model-agnostic, so strong engines get added once they are actually available.

What is FLUX-mimic?

FLUX-mimic is a video-action model built on FLUX 3 together with mimic robotics, one of the first early access partners. It uses the same video backbone to help robots predict the consequences of actions, and BFL cites deployments being tested with manufacturers like Audi. For ad makers it matters as proof that one model family now spans content generation and physical action.

Key Takeaways

  • Black Forest Labs announced FLUX 3 on July 23, 2026: one unified multimodal model that jointly learns from images, video, and audio, and extends to action prediction for robotics. It was announced, not launched.
  • FLUX 3 Video promises clips up to 20 seconds with native audio in a single generation, five input modes including video-to-video with character carry-over, and multi-shot chaining into minutes-long sequences.
  • Access is gated: FLUX 3 Video and Action are early access by application, FLUX 3 Image arrives in the coming weeks, general availability is 'to follow', there is no public price, and no aggregator like fal serves it yet.
  • The day-one performance numbers are BFL's own preliminary evaluations, including a 52% preference over Seedance 2.0 and Gemini Omni Flash; the methodology stays unpublished until broader availability, so treat them as claims, not benchmarks.
  • For ad makers the real promise is sound born with the picture and product consistency across motion. Until that reaches general availability, Novoads runs Seedance 2.0, Kling v3 Pro, Veo 3.1, and Sora today at roughly $2 to $11 per finished clip.
Mauricio Valdivia

Mauricio Valdivia

Founder of Novoads

Mauricio is the founder of Novoads, where he works to democratize video advertising with AI for brands in Latin America.

Ready to create video ads with AI?

Generate professional video ads in minutes, not weeks.

Start for $1