What Is LTX-2.3? Release Date, Paper, Specs and the LTX-2.5 Update
LTX-2.3 is Lightricks' open-weights audio-video model, released March 5, 2026. The paper, the specs, and what LTX-2.5 changed on Aug 11.
Mauricio Valdivia
·11 min

March 5, 2026: the video model you download, not rent
LTX announced LTX-2.3 in a press release dated March 5, 2026, with the open weights on Hugging Face and a local app, LTX Desktop, released alongside it. Five months later, on August 11, 2026, LTX launched LTX-2.5, and its own facts page now says "The latest model is LTX-2.5." Read this page as a dated reference.
The LTX-2.3 Release Date, in One Timeline
Most frontier video models reach you as a button that says "Get API key." You rent access by the second, the way Google's Veo 3.1 Lite bills it, and never hold the model. LTX-2.3 went the other way, and its launch has a clean paper trail.
The date comes from LTX's own press release, bylined "Zeev Farbman / March 5, 2026", which opens: "Today, LTX is launching LTX-2.3, a 20.9-billion parameter multimodal AI model." LTX's facts page, a page it publishes for AI assistants, agrees on the month: "LTX-2.3: Released March 2026."
Here is the family on one line, each date taken from an LTX surface or the arXiv record:
- October 23, 2025: LTX announces LTX-2, which its newsroom describes as "a foundation model with synchronized audio and video generation, native 4K at 50fps, and 10-second sequences."
- January 5, 2026: LTX open-sources LTX-2 and releases the full model weights.
- January 6, 2026: the LTX-2 paper is submitted to arXiv as 2601.03233.
- March 5, 2026: LTX-2.3 launches with open weights, together with LTX Desktop.
- August 11, 2026: LTX-2.5 launches. LTX's facts page labels it "The current release (August 2026)."
What shipped on March 5
The release bundled three things. First, the weights: "LTX-2.3 is being released as open weights, with model files publicly available on Hugging Face." Second, an app: "Alongside it, we are releasing LTX Desktop," which the release calls "free and open source under the Apache license." Third, a promise: "The developer CLI is coming soon."
Keep that Apache line attached to the app, which is where LTX put it. The model weights carry a different license, covered in the section on running it yourself.
Two parameter counts from one company
Size is where LTX's own pages disagree. The March 5 release says 20.9 billion parameters. The facts page, last updated August 2026, says "22 billion parameters", and the checkpoints on the Hugging Face card are named for 22B (ltx-2.3-22b-dev, for example). The same facts page puts the original LTX-2 at "Approximately 19 billion parameters."
The gap between 20.9 and 22 is small, and both figures come from the same company. If you cite one, name the page it came from. This post quotes 22B only where it quotes the facts page or a checkpoint name.

The LTX-2.3 Paper: What arXiv 2601.03233 Describes
The LTX-2.3 model card does not cite a paper of its own. It points back to the family's paper: "LTX-2 was presented in the paper" titled "LTX-2: Efficient Joint Audio-Visual Foundation Model", arXiv 2601.03233, submitted on January 6, 2026. LTX's facts page gives the same link when it answers "What is the research paper?"
So read the paper as the architecture 2.3 inherits, not as a changelog for 2.3. Three ideas in its abstract explain why one model produces sound and picture together.
Two streams, sized on purpose
The abstract says LTX-2 "consists of an asymmetric dual-stream transformer with a 14B-parameter video stream and a 5B-parameter audio stream." The authors describe the split as "allocating more capacity for video generation than audio generation."
That is a sensible trade. A second of video carries far more information than a second of sound, so the picture gets the bigger network.
Streams that listen to each other
The two streams are "coupled through bidirectional audio-video cross-attention layers." They are not run side by side and stapled together at the end. Each one attends to the other while it generates, so the audio is shaped by what the video is doing and the other way round.
That is the mechanism behind what the model card calls a model "designed to generate synchronized video and audio within a single model." For a UGC-style ad, where a believable talking person is the whole format, a voice that lands on the mouth without a separate lip-sync pass is the part that usually eats an afternoon.
A multilingual encoder and a guidance trick
The third idea is about your prompt. The authors write: "We employ a multilingual text encoder for broader prompt understanding and introduce a modality-aware classifier-free guidance (modality-CFG) mechanism for improved audiovisual alignment and controllability." For 2.3 specifically, LTX's repository says the checkpoints bundle the transformer and VAEs, "with the Gemma 3 text encoder downloaded separately."
The multilingual part is the quiet, ad-relevant bit: one creative concept often has to ship across markets, the problem we walk through in AI for advertising. The paper also claims that "the model achieves state-of-the-art audiovisual quality and prompt adherence among open-source systems." That is the authors' own evaluation of LTX-2, not an independent benchmark, and not a test of 2.3.
What Changed From LTX-2 to LTX-2.3
The Hugging Face card sums 2.3 up as "a significant update to the LTX-2 model with improved audio and visual quality as well as enhanced prompt adherence." The March 5 release and the facts page get more specific.
The five changes LTX lists
- A new VAE. The release describes "a new variational autoencoder (VAE) that preserves fine visual detail." The facts page calls it "a redesigned VAE producing sharper fine details and textures."
- A bigger text connector. The release says "a new text connector for improved prompt adherence"; the facts page says it is "4x larger."
- Native portrait video. The release lists "native portrait-mode support"; the facts page gives "native portrait video up to 1080x1920."
- Better image-to-video. The release promises "substantially improved image-to-video generation."
- Cleaner audio. "Audio generation has been refined through cleaner training data," per the release; the facts page adds "cleaner audio with fewer artifacts."
The facts page also lists "production-grade HDR output via IC-LoRA (Beta)." Every item above is LTX describing its own model. We have not measured any of them.
Why portrait is the change that matters for ads
Of the five, native portrait is the one that changes an ad workflow. Reels, TikTok and Shorts placements are vertical. A landscape-first model forces you to crop or reframe every clip, and a crop throws away the resolution you paid for. A model that renders 1080x1920 natively skips that step, which matters most when you are cutting TikTok ads in volume.
The checkpoints on the card
The 2.3 card publishes more than one model file, and picking the right one is half the job:
ltx-2.3-22b-dev: "The full model, flexible and trainable in bf16."ltx-2.3-22b-distilled: "The distilled version of the full model, 8 steps, CFG=1."ltx-2.3-22b-distilled-1.1: a later distilled build the card describes as "A different aesthetic experience and improved audio compared to v1.0." The card does not say when it was added.- Upscalers: spatial (x2 and x1.5) and temporal (x2) models for multi-stage pipelines.
Eight steps is the point of the distilled file. Use it while you are still finding the shot, and keep the full model for the final render.
Running LTX-2.3 Yourself: License and Hardware
Open weights are the headline, so be honest about both sides of the trade. Downloading a model is not the same as running it cheaply.
The license: open weights, not Apache
The model card lists the weights under the ltx-2-community-license-agreement. The March 5 release states the commercial line plainly: "The model weights are open and free to use; companies generating over $10 million in annual revenue require a commercial license."
The Apache license in that release belongs to LTX Desktop, the app, not to the weights. For a company past the revenue line, the weights need a commercial license from LTX, whatever license the app on top of them carries.
What owning the weights gets you
Holding the weights unlocks three things a rented endpoint rarely offers:
- Fine-tune the base model. The card says "The base (dev) model is fully trainable."
- Train your own adapters. LTX's repository ships a trainer package with "Training and fine-tuning tools for LoRA, full fine-tuning, and IC-LoRA", which is how a studio teaches the model a recurring spokesperson or a house look.
- Iterate without a meter. LTX's release pitches local runs "with no API calls, no per-generation fees, and no data leaving your machine."
For a studio with a signature style and an engineer to maintain it, that control is the whole reason to choose an open model. You are not buying clips. You are buying a model you can bend.
The hardware bill is the catch
The cost you avoid at the API moves to your own GPU and your own time. LTX's repository is full of ways to fit a big model onto a smaller machine. FP8 quantization "Enables lower memory footprint," and a block-streaming mode "Streams transformer blocks through the GPU one block at a time, so the full model runs on machines without enough memory to hold all its weights at once."
Those are good engineering answers. They are also a tell: running this well is an engineering project, not a signup.

What LTX-2.3 Can Do That Matters for Ads
Text-to-video is the part everyone expects. The more useful fact for ad teams is in the repository: "Every pipeline in this repository also runs on LTX-2.3." Four of those pipelines map onto jobs an ad team already does by hand:
- Dub a performance into a new script. The DubIt pipeline is described as "Dub-It: rephrasing while matching speaker identity and lip movements", and the LTX-2.3 models page lists a dedicated
LTX-2.3-22b-IC-LoRA-DubItadapter. Re-voicing one hero clip with new lines, without it reading as a bad overdub, is the same wedge that separates AI from a human UGC creator. - Start from the voiceover. The audio-to-video pipeline does "Audio-to-video generation conditioned on an input audio file", so a finished read can drive the picture.
- Fix one bad second. The retake pipeline can "Regenerate a specific time region of an existing video," so a near-perfect take with one flaw becomes a one-second fix, not a fresh roll of the dice.
- Call a camera move by name. The LTX-2.3 models page lists camera-control LoRAs for dolly in, dolly out, dolly left, dolly right, jib up, jib down and a static shot, though their file names carry the LTX-2 19B prefix.
Of the four, retake is the one that changes the economics. It turns a generation from a slot-machine pull into something closer to an editable asset, which is what a creative team actually needs from a tool.
LTX-2.3 vs LTX-2.5: Is 2.3 Still the Version to Use?
For new work, LTX's own answer is no. The LTX-2 repository now says "LTX-2.5 is the recommended model" and files 2.3 under a heading that reads "Legacy: LTX-2.3." There are still cases where 2.3 is the right pick.
What 2.5 adds, in LTX's words
LTX's facts page answers the comparison directly: "LTX-2.5 adds native multi-shot generation, a Video Editing IC-LoRA for editing real footage, a diffusion video decoder for sharper output, a custom Gemma 4 12B text encoder, a dedicated prompt enhancer, Auto Duration, a pretrained checkpoint for deep fine-tuning, cleaner licensing, and a substantially improved distilled model." That list is the vendor's description, not a test we ran. Our LTX-2.5 release breakdown separates what was new from what was inherited.
Here is how the two compare on the points LTX documents, with the API rows read on September 21, 2026:
| LTX-2.3 | LTX-2.5 | |
|---|---|---|
| Released | March 5, 2026 | August 11, 2026 |
| Text encoder | Gemma 3, downloaded separately | Gemma 4 12B, bundled |
| Final decoding | Redesigned VAE | Diffusion video decoder |
| API retake and extend | On ltx-2-3-pro | Not a listed model |
| API price, 1080p pro | $0.08 per second | $0.17 per second |
| LoRAs | Trained on 2.3 only | Trained on 2.5 only |
The price row comes from LTX's API pricing page, which lists 1080p text-to-video at $0.08 a second on ltx-2-3-pro and $0.17 on ltx-2-5-pro. The editing row comes from LTX's model docs, which say on 2.3 that "Audio-to-video, retake, extend, and reframe require ltx-2-3-pro", and from the same pricing page, whose Retake and Extend tables list ltx-2-3-pro as the only model.
Use 2.3 when, use 2.5 when
Match the version to the job, not to the version number:
- Reach for LTX-2.3 when you call LTX's hosted API and need retake, extend or reframe; when per-second API cost decides the budget; or when you already trained LoRAs on it, because the repository warns that "a LoRA only works with the model it was trained on."
- Reach for LTX-2.5 when you are starting fresh; when you want several connected shots from one generation, which LTX describes as "Native Multishot: A single generation produces multiple connected shots"; or when you want the new decoder, which LTX says "produces sharper faces in close-up, more legible text and signage, and fewer smears in fast motion."
Our read: 2.3 is now a maintenance choice. Keep it where your pipeline already depends on it, and start anything new on 2.5.

How Novoads Solves the Same Job Without the Setup
Novoads does not run LTX, in either version. It runs managed video models behind a no-setup workflow: Seedance 2.5, Seedance 2.0, Seedance 2.0 Mini, Kling v3 Pro, Google Veo 3.1 and Omni Flash. There is no GPU to provision and no pipeline to wire. In Novoads the flow skips the infrastructure:
- Upload a product photo and write or auto-generate a script.
- Pick an AI actor from more than 100, or create a custom actor from a photo of someone holding the product when it has to be in hand.
- Generate a vertical ad with voice, lip-sync and captions in about four minutes.
- Localize the same spot with voices in 31 languages.
The cost is per clip instead of per GPU: from about 25 cents for a five-second Seedance 2.0 Mini clip to about $4 for an eight-second Google Veo 3.1 clip, and about $10 for a full 30-second Seedance 2.5 take. Novoads starts at $15/month (Starter, 200 credits per month). Plus is $49/month (1,000 credits), and Ultra starts at $129/month (3,000 credits) for more volume. All plans are published on novoads.ai/pricing.
The choice is about operating model, not which model is smarter. Reach for LTX when you have GPUs, engineers and a reason to own the weights. Reach for a managed model when you need a finished ad this week, the same way you would compare the AI video ad platforms before betting a quarter on one. If the UGC format itself is new to you, start with what a UGC creator is.
The weights are open. The bottleneck moved.
LTX-2.3 was a real milestone, but it is easy to misread what kind. It did not make a finished ad cheaper to ship for most teams, because the API bill it removed came back as a GPU bill and an engineering project. What it did was hand the model itself, weights, audio and control, to the people who want to own their stack rather than rent it from an endpoint that can be retired, as OpenAI's video API shutdown shows. LTX-2.5 continued the same bet five months later. The open question was never whether the weights would open. It was where the work would go once they did, and the answer has not changed between versions: it moved from the render to the rig.
Frequently Asked Questions
When was LTX-2.3 released?
On March 5, 2026. LTX's press release of that date announces the launch, says the model weights are on Hugging Face, and releases a local app called LTX Desktop alongside the model. LTX's facts page gives the same month: 'Released March 2026.'
Is there an LTX-2.3 paper?
The paper that describes the model family is 'LTX-2: Efficient Joint Audio-Visual Foundation Model', arXiv 2601.03233, submitted on January 6, 2026. The LTX-2.3 model card links that paper rather than one of its own, and LTX's facts page gives the same link when it answers 'What is the research paper?'
Is LTX-2.3 the latest LTX model?
No. LTX launched LTX-2.5 on August 11, 2026, and LTX's facts page says 'The latest model is LTX-2.5.' LTX-2.3 is the March 2026 release, and LTX's code repository now lists it under Legacy.
What is the difference between LTX-2 and LTX-2.3?
LTX lists a new VAE for finer detail, a larger text connector for prompt adherence, native portrait video up to 1080x1920, improved image-to-video and cleaner audio. On size, LTX's facts page puts LTX-2 at approximately 19 billion parameters and LTX-2.3 at 22 billion, while the March 5 release gives 20.9 billion for 2.3.
Can I use LTX-2.3 commercially?
LTX's March 5 release says the model weights are open and free to use, and that companies generating over $10 million in annual revenue require a commercial license. The model card lists the weights under the LTX-2 Community License Agreement. The Apache license in the same release covers the LTX Desktop app, not the weights. Running the model still costs GPU time and engineering.
Can I use LTX-2.3 in Novoads?
No. Novoads does not run LTX-2.3 or LTX-2.5. It runs managed video models with no setup (Seedance 2.5, Seedance 2.0, Seedance 2.0 Mini, Kling v3 Pro, Google Veo 3.1 and Omni Flash), with clips running from about 25 cents for a five-second Seedance 2.0 Mini clip to about $4 for an eight-second Google Veo 3.1 clip, and about $10 for a full 30-second Seedance 2.5 take.
Key Takeaways
- LTX-2.3 launched on March 5, 2026, the date on LTX's own press release, with the open weights on Hugging Face and a local app, LTX Desktop, released alongside it.
- The paper to cite is the LTX-2 paper, arXiv 2601.03233, submitted on January 6, 2026, which the 2.3 model card links. It describes a 14B-parameter video stream and a 5B-parameter audio stream joined by cross-attention.
- LTX lists five changes over LTX-2: a new VAE, a larger text connector, native portrait video, better image-to-video and cleaner audio. The size depends on the page: the March release says 20.9 billion parameters, LTX's facts page says 22 billion.
- LTX-2.3 is no longer LTX's latest model. LTX-2.5 launched on August 11, 2026, and LTX's own code repository now calls 2.5 the recommended model and files 2.3 under Legacy.
- 2.3 still has a place: as of September 21, 2026, LTX's API docs list retake, extend and reframe only for ltx-2-3-pro, and a LoRA trained on 2.3 does not carry over to 2.5. Novoads runs neither version.
Sources
- •LTX: LTX 2.3 Offers Creative Intelligence Without the Cloud Costs (press release, March 5, 2026)
- •LTX: Official AI information (facts page, last updated August 2026)
- •LTX-2: Efficient Joint Audio-Visual Foundation Model (arXiv 2601.03233)
- •Lightricks: LTX-2.3 model card (Hugging Face)
- •Lightricks: LTX-2 code repository (GitHub)
- •LTX API docs: LTX-2.3 model page




