LTX-2: Efficient Joint Audio-Visual Foundation Model
This work addresses the prevailing limitation of existing text-to-video diffusion models in generating high-quality, synchronized audio that aligns semantically, emotionally, and atmospherically with the visual content. To this end, we propose a unified audio-visual generative foundation model featuring an asymmetric dual-stream Transformer architecture—comprising a 14B-parameter video stream and a 5B-parameter audio stream. The model leverages modality-aware classifier-free guidance (CFG), cross-modal AdaLN, and bidirectional audio-visual cross-attention mechanisms to achieve efficient co-generation and precise temporal alignment. Integrated with temporal positional encoding and a multilingual text encoder, our approach achieves state-of-the-art audio-visual quality and prompt fidelity within an open-source framework, matching the performance of closed-source counterparts while significantly reducing computational overhead and inference latency.