Skip to main content

A 22B Open-Weights Video Model Now Renders Multishot Ads on Hardware You Already Own

LTX shipped a new open-weights release on August 11, 2026, and the genuinely new capability is multishot: one generation that holds character, scene and voice across cuts. Here is what changed, what was already there, and what it costs.

Mauricio Valdivia

Mauricio Valdivia

·11 min

A 22B Open-Weights Video Model Now Renders Multishot Ads on Hardware You Already Own

The headline feature is the cut, not the resolution

A skincare brand books a fifteen-second spot as three shots: a hand lifting the jar off a shelf, a face two seconds after the first application, the label held up to camera. Three shots, one person, one room, one voice. On most AI video models that is three separate renders, and then an afternoon in an editor matching skin tone and window light across the joins so the cuts do not announce themselves.

That afternoon is what LTX aimed at on August 11, 2026. The company released a 22-billion-parameter open-weights world model, and its model card describes native multishot generation as connected scenes produced "in a single pass: multiple shots that hold character identity, environment, lighting, voice, and visual style across cuts". Then it adds the parenthetical that dates the change precisely: "previous versions produced a single continuous shot."

That is the news. It is worth saying out loud because most of the coverage you will read leads on synchronized audio, open weights and 4K instead, and all three of those are older than this release. This post separates the two piles. What shipped on launch day, what was already there, what the thing costs, and which of those facts survive contact with a media buyer's calendar.

What actually shipped on launch day

LTX did not describe this as an incremental patch. The launch post says the team "rebuilt nearly every stage of the generation pipeline rather than bolting capabilities onto an older core", and the published component list backs that up: the weights arrive as separate files per stage rather than one bundle. Seven changes carry names in the vendor's own launch list:

  • Native multishot generation, rendering a full sequence as one output that holds character, scene and voice across cuts
  • A new diffusion video decoder that replaces the previous reconstruction stage and cuts artifacts in high motion
  • A custom Gemma 4 language backbone plus a dedicated prompt enhancer for complex, multi-subject prompts
  • Diffusion Fidelity Rendering, which builds motion in a compressed latent space and anchors detail with adaptive keyframes
  • A pretrained physical-AI checkpoint tuned for robotics teams to fine-tune on their own domain data
  • A substantially better distilled model, aimed at the same quality at lower cost and faster inference
  • Local inference work with NVIDIA for RTX GPUs and DGX Spark, cutting memory requirements

Three of those matter to anyone making ads. The rest are infrastructure for a different audience, and it is worth knowing which is which before a vendor quotes you on robotics benchmarks.

Multishot, and the parenthetical that dates it

The launch post's own phrasing is that "native multishot generation renders a full sequence as one output, holding character, scene, and voice across cuts". The API documentation says the same thing in engineering terms, and enumerates exactly what is supposed to survive the cut:

  • Character identity, so the face in shot three is the face from shot one
  • Environment, so the room does not quietly redecorate itself
  • Lighting, the one continuity error a colourist can usually rescue
  • Visual style, including lens character and grade
  • Voice, the one nobody can rescue

For anyone who has assembled an ad out of clips, the interesting item there is not the picture. It is voice. Visual continuity across a cut is a colour problem you can mostly fix downstream. Vocal continuity is not: a narrator whose timbre shifts at the two-second mark reads as two different people, and no grade rescues it. Holding voice across a cut inside one generation is the part that removes editing work rather than relocating it.

A new decoder, and where the artifacts went

The second named change is a new diffusion video decoder. The launch post frames it as reducing "visual artifacts in high motion while keeping the high compression ratio of LTX", and the model card is blunter about what it displaced: it "replaces the VAE reconstruction stage", with better faces, textures and on-screen text.

On-screen text is the quiet one there. If you have ever generated a product shot and watched the label melt into approximate lettering, you know that reconstruction, not the diffusion process, is usually where a legible logo goes to die. We wrote about the same failure surfacing on other engines in our notes on keeping an AI actor consistent across ad variations.

A language backbone and a prompt enhancer

The third change is upstream of the picture entirely. LTX says a "custom Gemma 4 language backbone and dedicated prompt enhancer read and understand complex, multi-subject prompts more accurately", and the checkpoint list confirms a 12B text encoder shipping as its own file. The model card explains the failure it targets: holding multiple characters, camera moves, lighting and actions together "instead of dropping details across a longer sequence".

That is a direct consequence of multishot. A three-shot prompt is roughly three times the instruction load of a single-take prompt, and a text encoder that drops the third clause gives you a beautiful first shot and two strangers. If you write prompts for a living, our camera movement prompt guide covers the same load problem from the authoring side.

A real UGC creator filming a product testimonial on a phone
Novoads · UGC video ads with AI, ready in minutes.
Try now

The three features that are not news

Here is where the launch coverage and the launch diverge. Three capabilities are being reported as new that the vendor itself does not list as new, and knowing which is which changes how you brief a team.

  • Synchronized audio and video dates to the previous foundation release, announced October 23, 2025
  • High-resolution output came from that same announcement, and is now a variant-level decision rather than a model-level one
  • Open weights is the company's standing posture: the same newsroom archive carries a March 5, 2026 card describing the previous model shipping "with local inference and open weights"

Synchronized audio has been the pitch since 2025

The newsroom page that carries the announcement also carries an archive card for the previous foundation release, dated October 23, 2025, describing "synchronized audio and video generation, native 4K at 50fps, and 10-second sequences in a single production-ready system". That is the same pitch, roughly ten months earlier, sitting in the sidebar of the page everyone is citing.

The current model card confirms the capability is present, describing the established use as "generating synchronized, high-fidelity video and audio from text, image, and video inputs". Present, though, is not new. Audio appears in this release's feature list as voice continuity inside multishot, which is a genuinely different claim.

Where the resolution claim actually lives

The same archive card is where native 4K comes from. In this release, resolution is a variant decision rather than a model property, which we will get to in a moment. Repeating a flat "native 4K" headline attributes to the whole model something only half of it does.

Why the distinction changes your brief

This is not pedantry, it is procurement. The two briefs buy different things.

  • "The new model with synchronized audio" describes the state of the art in October 2025. A partner will quote it as a normal generation job, with an editor assembling the cuts, because that is what it has always meant.
  • "The model that holds a character and a voice across three cuts in one render" describes a change to the edit itself. The estimate should shrink by the assembly hours it deletes, and if it does not, ask why.

The general habit is worth keeping past this release: read the vendor's own "what's new" list before you read anyone's summary of it. It is the cheapest fact-check in this category, and the same discipline is what we applied when Seedance announced its 30-second single take.

One further caution, and it applies to every number you will see attached to this launch. As of publication, the quality and speed comparisons circulating are LTX's own. The company publishes a comparison chart on its model page and labels it "Preliminary results, expected to evolve as evaluation expands", noting the clips were "graded by automated scoring rather than human viewers". Its speed chart credits a third-party host for the API figures and its own hardware for the local ones. None of that is dishonest, and all of it is disclosed. It is simply not independent testing, which has not landed yet. On a same-day launch that is the normal state of the world, and the correct response is to treat the numbers as claims rather than findings and to run your own comparison on your own creative.

Fast and pro are not the same product

The single most useful thing in LTX's documentation is also the least dramatic: a support matrix showing that the two variants are not a quality dial on one product. They have different resolution ranges and different duration options, and choosing wrong costs you a re-render.

The resolution split

The documentation is explicit. "ltx-2-5-pro generates at 720p and 1080p only", and for 1440p or 4K you are sent to the fast variant. So the accurate sentence for an ad brief is up to 4K on the fast variant, never a flat native-4K claim about the model. The docs list 4K as 3840x2160 in landscape and 2160x3840 in portrait, which is the orientation that matters for paid social.

Duration is a function of resolution

The matrix pairs each resolution with the clip lengths available to it, and the pattern is consistent across both variants: the higher you push resolution or frame rate, the shorter the available clip. LTX's hosted API documents a text-to-video request as "up to 4K resolution and 20 seconds per request", and the per-variant table is where that ceiling gets spent.

Decision pointFast variantPro variant
Top resolutionUp to 4K1080p
1440p availableYesNo
Longest clipsAt lower resolutionsAt lower resolutions
Shared across bothText, image and audio to video · 16:9 and 9:16 · optional silent output · automatic duration

Which one a campaign wants is usually obvious once the placement is named.

Use the fast variant when:

  • The deliverable is a 9:16 paid-social cut the platform will re-compress anyway
  • You are iterating on hooks and want more attempts per dollar
  • Someone has genuinely asked for a 4K master for connected TV or an in-store screen

Use the pro variant when:

  • The shot is the hero and the product surface is glossy or textured enough to punish a weak render
  • The cut is going into a brand film where a 1080p master is the accepted deliverable
  • Fidelity per frame matters more to the client than a resolution number on paper

The same trade shows up across engines, and we walked through it comparing Kling and Seedance on ad work.

Automatic duration, and the credit hold nobody expects

The release added automatic duration: send a null duration and the model picks the clip length from the prompt. It is the sort of feature that reads as convenience and behaves as an accounting event, because the length is not known until the job finishes.

LTX documents the consequence plainly. On a prepaid account, credits are held against the longest duration your resolution and frame rate allow, so "a request that would have produced 6 seconds is still declined if you cannot cover 20 on" the fast variant. Postpaid accounts hold nothing. There is one more edge worth writing on a sticky note: automatic duration cannot be combined with a last-frame image input, because a fixed final frame requires a known length.

The Novoads app: pick an AI actor, write a script, generate a UGC ad
Novoads · UGC video ads with AI, ready in minutes.
Try now

What it costs, and who is quoting the number

Two organizations publish per-second rates for this model. They are separate ladders on separate meters, and merging them into one number is how a media plan ends up wrong by a third.

LTX's own API rates

LTX's documentation states that video generation "is billed per second of output video". Its published table for this release reads:

  • Fast variant, 720p: $0.09 per second
  • Fast variant, 1080p: $0.13 per second
  • Fast variant, 1440p: $0.19 per second
  • Fast variant, 4K: $0.30 per second
  • Pro variant, 720p: $0.12 per second
  • Pro variant, 1080p: $0.17 per second

One meter behaves differently from the rest: audio-to-video is billed on the duration of the input audio rather than the output, so a long voice track is priced before you have seen a single frame.

fal's rates for the same endpoints

The model is also hosted on fal, which publishes its own prices. On the pro text-to-video endpoint fal states your request "will cost $0.12 per second for 720p or $0.17 per second for 1080p", and on the fast endpoint it quotes $0.09 per second at 720p, adding that "native audio is included at every resolution". Those happen to line up with LTX's own ladder today, which is a fact about this week and not a rule. When two meters for one model do diverge, they can diverge hard: the same avatar model lists 71% higher per second on a reseller than from the vendor that built it. Quote the meter you are actually billed on, name it in the plan, and re-check it before the invoice.

A worked example: a three-shot ad at 1080p

Take the skincare spot from the top. Fifteen seconds, three shots, pro fidelity at 1080p, on LTX's own published rate.

  • One clean render: 15 seconds at $0.17 per second is $2.55
  • Six attempts before the third shot behaves: $15.30 of compute for one finished cut
  • Twenty variations of that spot for a test: $306 if every one takes six attempts
  • What that replaces: three separate renders per variation plus the editing pass that matches them

The number to compare against is not a cheaper per-second rate. It is the afternoon of editing that multishot is supposed to delete, plus the re-renders you would otherwise spend matching three separately generated clips. That is the same arithmetic we ran when we measured real render times across models, and the conclusion held there too: the per-second price is rarely the expensive part of a creative cycle. Per-clip pricing on other engines lands in a similar band, as our breakdown of Seedance pricing shows.

What open weights gets you, in files

The launch calls this "the most capable open weights world model on the market", and the reason LTX gives for working this way is that open weights let teams own their hardware, their intellectual property and their model. Both halves of that are true, and neither is quite the same as open source.

The license, in the vendor's own words

The Hugging Face repository carries the LTX-2.x Community License. The model card offers "Commercial and production use at no cost under the LTX-2.x Community License" for organizations under a stated revenue threshold, notes that transfer of fine-tunes may require a paid license, and adds that revenue "is measured across the whole entity, including subsidiaries and affiliates under common control". The launch post phrases the threshold as organizations under $10M in annual recurring revenue.

Three cautions, all of them from LTX's own pages rather than anyone's commentary.

  • The summary is not the contract. The card says the full binding terms live in the license file, so read the file before you build a business on the headline phrasing.
  • The download is gated. Hugging Face requires you to agree to share your contact information to access the model, so no-cost use is not the same as anonymous acquisition.
  • The threshold counts your whole group. Revenue is measured across the entity including subsidiaries and affiliates under common control, which is the clause an agency holding company should read twice.
A UGC creator filming a product review without a film crew
Novoads · UGC video ads with AI, ready in minutes.
Try now

Eight files, not one

The weights ship as "a split, comfy-aligned pack (one .safetensors per component) rather than a single monolith". The components are worth listing, because a half-downloaded pipeline fails in confusing ways.

  • Distilled transformer and a full, trainable transformer, both at 22B
  • Quantized transformer variants for ComfyUI and for Blackwell-class hardware
  • A 12B text encoder shipping as its own file
  • Two video autoencoders, one higher quality and heavier, one faster and lighter
  • An audio autoencoder with vocoder, which is where the sound actually comes from
  • A duration head, the patch that powers automatic clip length
  • A distilled LoRA for the trainable-transformer workflows

There is one more detail worth budgeting for: the spatial and temporal upscalers are "still required for multi-stage" and are not in this repository, so a full pipeline still pulls a file from the previous release.

The hardware, and what is still missing

LTX says it worked with NVIDIA to optimize the model "for local inference on NVIDIA RTX GPUs and DGX Spark, cutting memory requirements", and the launch shipped with ComfyUI as a day-one partner, so the workflow templates exist on arrival. The inference package added support in its v1.2.0 release, published the same day. Existing adapters mostly carry over: LTX reports that "the large majority of LoRAs and IC-LoRAs trained on" the previous release run on this one without changes, with a small number of exceptions to validate before production.

What is not there yet is worth stating too, and LTX states it itself.

  • Retake, extend and reframe are marked unavailable on both variants in the support table as of launch day
  • The pricing page lists those editing endpoints only for the previous release
  • Precise video editing is described on the model page as being in beta

If your workflow depends on regenerating a section of a finished clip rather than re-rolling the whole thing, that path still runs on the older model. Teams with a standing variation pipeline should read that as a sequencing note rather than a blocker, which is the sort of dependency our creative operations guide exists to catch before it reaches a calendar.

How Novoads solves continuity when you are not running your own GPUs

Novoads does not run LTX models. Our picker offers Seedance, Kling, Sora and Veo for video, plus GPT Image 2 and Nano Banana Pro for stills, and this release is an adjacent development rather than something you can click on our platform today. Saying so plainly matters, because the interesting question this launch raises is not which logo renders your frames.

Multishot solves continuity inside one generation, which is the right answer if you own the GPUs and the pipeline. Most performance teams own neither. Their continuity problem lives one level up: the same actor, the same product and the same script surviving across twenty ad variations that will be generated separately over three weeks and tested against each other. Novoads holds that layer. You upload a product image or write a script, pick an AI actor, and get an ad-ready vertical or horizontal file, with the actor and script held constant while the angle changes. You can try it for $1, which covers three days of access, then $49 per month. If you are still deciding what to vary, our guide to ad creative testing is the better first read.

Continuity is the spec that decides the edit

Every generation of AI video models has had one number the marketing leans on and one property that actually changes the work. For two years the number was resolution and the property was duration. This release suggests the next pair: the number is still resolution, and the property is now whether the model can hold a person across a cut.

That is a better spec to shop on than pixels, because it maps to the only question a creative team asks about a tool: how much of the finished ad comes out of the machine, and how much still has to be assembled by a human matching light. Multishot is the first honest attempt at moving that line. Watch what the vendors put on the voice row of their spec sheets next, because that is where the remaining afternoons live.

Frequently Asked Questions

What is actually new in this release?

Native multishot generation is the headline: one generation that produces several connected shots holding character, scene, lighting and voice across cuts. Alongside it LTX shipped a new diffusion video decoder that replaces the VAE reconstruction stage, a custom Gemma 4 language backbone with a dedicated prompt enhancer, a rendering technique the company calls Diffusion Fidelity Rendering, a pretrained checkpoint for physical AI and robotics fine-tuning, a substantially better distilled model, and local inference work done with NVIDIA for RTX GPUs and DGX Spark.

Is synchronized audio new here?

No. LTX's own newsroom archive dates synchronized audio and video generation to the 2.x foundation launch on October 23, 2025. The current model card confirms the model generates synchronized video and audio from text, image and video inputs, but that is inherited behaviour rather than this release's news. Any coverage leading on it is describing a feature that is roughly ten months old.

Does it really do 4K?

On one variant. LTX's documentation states that the pro variant generates at 720p and 1080p only, and directs you to the fast variant for 1440p or 4K. So the honest phrasing is up to 4K on the fast variant. The hosted API also documents a per-request ceiling for text-to-video of up to 4K resolution and 20 seconds.

What does it cost per second?

It depends whose meter you are on, and the two ladders should never be merged. LTX's own API pricing page lists the fast variant at $0.09 per second for 720p, $0.13 for 1080p, $0.19 for 1440p and $0.30 for 4K, with the pro variant at $0.12 and $0.17. The hosting platform fal publishes its own rates for the same endpoints, quoting $0.12 per second for 720p and $0.17 per second for 1080p on pro.

Is this open source?

It is open weights, which is not the same thing. The Hugging Face repository carries the LTX-2.x Community License and gates the download behind agreeing to share your contact information. The model card describes commercial and production use at no cost under that license for organizations under a revenue threshold, notes that transfer of fine-tunes may require a paid license, and points to the license file itself for the binding terms.

Can I generate ads with this model inside Novoads?

No. Novoads does not run LTX models. The engines in the picker are Seedance, Kling, Sora and Veo, plus GPT Image 2 and Nano Banana Pro for stills. This release matters to ad teams as an adjacent development in how continuity across cuts gets solved, not as something you can click on our platform today.

Key Takeaways

  • The genuinely new capability in this release is native multishot: LTX's model card describes connected scenes generated in a single pass, and adds that previous versions produced a single continuous shot. That parenthetical is the whole story.
  • Synchronized audio and high-resolution output are not news. LTX's own newsroom archive dates them to the 2.x launch on October 23, 2025. Treat them as inherited table stakes, not as this release's headline.
  • Resolution is a variant decision, not a model decision. LTX's docs state that the pro variant generates at 720p and 1080p only, and send you to the fast variant for 1440p or 4K.
  • Two organizations publish per-second rates for this model, on two separate meters that can diverge. LTX's own API docs and fal both quote $0.12 per second at 720p on the pro endpoint today, and LTX's fast variant starts at $0.09. Always name whose price you are quoting.
  • Open weights here means the LTX-2.x Community License plus a gated download, not OSI open source. The model card puts commercial use at no cost under that license and points to the license file for the binding terms.
Mauricio Valdivia

Mauricio Valdivia

Founder of Novoads

Mauricio is the founder of Novoads, where he works to democratize video advertising with AI for brands in Latin America.