LTX-2.5: 22B Open Audio-Video Model with Native Multishot and 4K HDR
LTX-2.5 is Lightricks' 22B open-weights audio-video model with native multishot scenes, auto duration and 4K HDR output, supported in ComfyUI on day one.
LTX-2.5
VideoAudio-VideoText-to-VideoImage-to-VideoDiTLightricksLightricks' 22B-parameter open-weights audio-video foundation model, built on an asymmetric dual-stream diffusion transformer that generates video and audio jointly. It adds native multishot scenes, auto duration, 4K HDR output and Diffusion Fidelity Rendering, and ships as both a full dev checkpoint and a distilled few-step variant.
| Developer | Lightricks |
| Release Date | 2026-08 (open weights) |
| Architecture | Asymmetric dual-stream DiT (22B), joint video and audio generation |
| Text Encoder | Gemma 4 12B with learned projection |
| License | LTX-2.x Community License |
| Generation Modes | Text-to-video, image-to-video, video-to-video, text-to-audio, audio-to-video |
| Quantized Weights | INT8 convrot, NVFP4 (distilled transformer) |
What is LTX-2.5?
LTX-2.5 is the successor to LTX-2.3 in Lightricks' open video family. It is a 22B-parameter asymmetric dual-stream diffusion transformer that generates video and audio jointly through bidirectional cross-attention with modality-aware classifier-free guidance, and it pairs a Gemma 4 12B text encoder with a learned projection.
The release ships as a full dev checkpoint and a distilled few-step variant, and covers text-to-video, image-to-video, video-to-video, text-to-audio and audio-to-video generation.
What is new in LTX-2.5?
- Native multishot. Connected scenes, wide, medium and close-up, are generated in a single pass while character, environment, lighting and voice stay consistent across the cuts.
- Auto duration. Clip length is predicted from the action being described instead of being set by hand as a frame count.
- Native 4K HDR. The model targets high-resolution, high-dynamic-range output with HDR pipelines built into the official workflows.
- Diffusion Fidelity Rendering (DFR). A rendering pass aimed at real footage, editable results and cinema-grade EXR output for production pipelines.
Weights and variants
The official Hugging Face repository provides the full weight set: the 22B dev transformer, the distilled transformer, ComfyUI-ready INT8 convrot and NVFP4 quantizations of the distilled build, the Gemma 4 12B text encoder (BF16 and INT8 convrot), separate video and audio VAEs, 2x latent spatial and temporal upscalers, a distilled LoRA and a duration-head model patch.
Community adapters for LTX-2.5, such as the CQ enhancer LoRAs and the Ripple and AInVFX Fluid IC-LoRAs, are curated in the Model Hub downloads for this version.
ComfyUI support
LTX-2.5 is supported in ComfyUI on day one, with official workflow templates for text-to-video and image-to-video in the built-in template gallery:
The official ComfyUI-LTXVideo wrapper carries a matching 2.5 example set covering single-stage and two-stage upscaled T2V/I2V, IC-LoRA inpainting and outpainting, motion tracking, union control and text-to-audio.
Guides and workflows related to this model series.
Comments
Sign in with GitHub to join the discussion.