LTX-2.5 is a 22-billion-parameter open-weight world model that generates synchronised video and audio from text, images or existing clips, with support for native 4K HDR output.
LTX-2.5 is an open-weight world model developed by LTX, a company spun out of Lightricks. The model is designed for video generation, real-time applications and physical AI use cases such as robotics simulation. It can generate text-to-video, image-to-video and video-to-video content with synchronised audio, supporting resolutions from 720p to native 4K and clip lengths of six to 20 seconds.
The architecture is based on a 22-billion-parameter asymmetric dual-stream diffusion transformer with independent streams for video and audio linked through bidirectional cross-attention. This enables both forms of media to be generated together instead of producing the video and audio separately. It also uses a 12-billion-parameter Gemma 4 text encoder and prompt enhancer to interpret instructions involving characters, camera movements, lighting and actions.
One of the aims of LTX-2.5 is to reduce the computing requirements for local video generation. The distilled checkpoint uses an eight-step generation process, while FP8 quantisation can reduce VRAM requirements by up to 40%. Local deployment is supported through ComfyUI and Diffusers, with hardware requirements starting at 16 GB of VRAM depending on the checkpoint and quantisation used.
The model features native multishot generation, where several consecutive shots are generated while preserving character identity, lighting, environment, sound and visual style within a single generation. It also uses Diffusion Fidelity Rendering, which allocates rendering effort according to scene complexity. A new video decoder is designed to improve fine details and stability in high-motion scenes.
LTX-2.5 is available as open weights on Hugging Face and is integrated into ComfyUI. Developers can also use the LTX API or a local Python-based pipeline. The model can be fine-tuned using the LTX-2 Trainer, allowing developers to adapt it to their own datasets and applications.
















































































