LTX-2.5: 22B Open Audio-Video Model with Native Multishot and 4K HDR

ComfyUI Wiki

LTX-2.5 is Lightricks' 22B open-weights audio-video model with native multishot scenes, auto duration and 4K HDR output, supported in ComfyUI on day one.

L

LTX-2.5

VideoAudio-VideoText-to-VideoImage-to-VideoDiTLightricks

Lightricks' 22B-parameter open-weights audio-video foundation model, built on an asymmetric dual-stream diffusion transformer that generates video and audio jointly. It adds native multishot scenes, auto duration, 4K HDR output and Diffusion Fidelity Rendering, and ships as both a full dev checkpoint and a distilled few-step variant.

DeveloperLightricks
Release Date2026-08 (open weights)
ArchitectureAsymmetric dual-stream DiT (22B), joint video and audio generation
Text EncoderGemma 4 12B with learned projection
LicenseLTX-2.x Community License
Generation ModesText-to-video, image-to-video, video-to-video, text-to-audio, audio-to-video
Quantized WeightsINT8 convrot, NVFP4 (distilled transformer)

What is LTX-2.5?

LTX-2.5 is the successor to LTX-2.3 in Lightricks' open video family. It is a 22B-parameter asymmetric dual-stream diffusion transformer that generates video and audio jointly through bidirectional cross-attention with modality-aware classifier-free guidance, and it pairs a Gemma 4 12B text encoder with a learned projection.

The release ships as a full dev checkpoint and a distilled few-step variant, and covers text-to-video, image-to-video, video-to-video, text-to-audio and audio-to-video generation.

What is new in LTX-2.5?

  • Native multishot. Connected scenes, wide, medium and close-up, are generated in a single pass while character, environment, lighting and voice stay consistent across the cuts.
  • Auto duration. Clip length is predicted from the action being described instead of being set by hand as a frame count.
  • Native 4K HDR. The model targets high-resolution, high-dynamic-range output with HDR pipelines built into the official workflows.
  • Diffusion Fidelity Rendering (DFR). A rendering pass aimed at real footage, editable results and cinema-grade EXR output for production pipelines.

Weights and variants

The official Hugging Face repository provides the full weight set: the 22B dev transformer, the distilled transformer, ComfyUI-ready INT8 convrot and NVFP4 quantizations of the distilled build, the Gemma 4 12B text encoder (BF16 and INT8 convrot), separate video and audio VAEs, 2x latent spatial and temporal upscalers, a distilled LoRA and a duration-head model patch.

Community adapters for LTX-2.5, such as the CQ enhancer LoRAs and the Ripple and AInVFX Fluid IC-LoRAs, are curated in the Model Hub downloads for this version.

ComfyUI support

LTX-2.5 is supported in ComfyUI on day one, with official workflow templates for text-to-video and image-to-video in the built-in template gallery:

The official ComfyUI-LTXVideo wrapper carries a matching 2.5 example set covering single-stage and two-stage upscaled T2V/I2V, IC-LoRA inpainting and outpainting, motion tracking, union control and text-to-audio.

Guides and workflows related to this model series.

No articles found.

Comments

Sign in with GitHub to join the discussion.

Loading comments…