Institution profile

Lightricks

Industry researcheurope · il
Official website
Research library3linked papers
Opportunities0open roles
Selected work

Representative Papers

LTX-2: Efficient Joint Audio-Visual Foundation Model

Jan 06, 2026arXiv.org

This work addresses the prevailing limitation of existing text-to-video diffusion models in generating high-quality, synchronized audio that aligns semantically, emotionally, and atmospherically with the visual content. To this end, we propose a unified audio-visual generative foundation model featuring an asymmetric dual-stream Transformer architecture—comprising a 14B-parameter video stream and a 5B-parameter audio stream. The model leverages modality-aware classifier-free guidance (CFG), cross-modal AdaLN, and bidirectional audio-visual cross-attention mechanisms to achieve efficient co-generation and precise temporal alignment. Integrated with temporal positional encoding and a multilingual text encoder, our approach achieves state-of-the-art audio-visual quality and prompt fidelity within an open-source framework, matching the performance of closed-source counterparts while significantly reducing computational overhead and inference latency.

6 citationsRead paper

AVControl: Efficient Framework for Training Audio-Visual Controls

Mar 25, 2026

Existing audio-visual generation methods struggle to efficiently support diverse control signals, often relying on fixed models or requiring costly architectural modifications. This work proposes a lightweight and extensible framework built upon the LTX-2 joint audio-visual foundation model, where each control modality—such as depth, pose, camera trajectory, sparse motion, and audio—is trained as an independent LoRA module. These modules are integrated into the backbone network via a parallel canvas mechanism that injects them as additional attention tokens, eliminating the need for any architectural changes. This approach enables modular audio-visual control for the first time, allowing flexible composition of control signals and efficient training. On the VACE benchmark, the method outperforms existing approaches in depth- and pose-guided generation, image inpainting, and extrapolation tasks, achieves competitive performance in camera control and audio-visual synthesis, and significantly reduces training costs.

0 citationsRead paper

Enhancing Forecasting with a 2D Time Series Approach for Cohort-Based Data

Aug 21, 2025

This work addresses the challenge of cohort-based two-dimensional time series forecasting under few-shot learning conditions. Methodologically, it introduces a novel neural network framework that jointly models cohort-level structural patterns and temporal dynamics: (1) cohort embedding is innovatively integrated into 2D time series modeling to co-learn intra-cohort temporal evolution (across time) and inter-cohort heterogeneity (across cohorts); (2) a lightweight spatiotemporal attention mechanism is designed to enhance generalization under sparse data regimes. Extensive experiments on multi-source real-world datasets from finance and marketing domains demonstrate that the proposed model achieves average MAE improvements of 12.7%–23.4% over state-of-the-art baselines—including TCN, Transformer, and matrix factorization–based approaches—while significantly improving both predictive accuracy and business interpretability in low-data scenarios. The architecture exhibits strong practical deployability due to its parameter efficiency and robustness.

0 citationsRead paper
Recent publications

Latest Papers

AVControl: Efficient Framework for Training Audio-Visual Controls

Mar 25, 2026

Existing audio-visual generation methods struggle to efficiently support diverse control signals, often relying on fixed models or requiring costly architectural modifications. This work proposes a lightweight and extensible framework built upon the LTX-2 joint audio-visual foundation model, where each control modality—such as depth, pose, camera trajectory, sparse motion, and audio—is trained as an independent LoRA module. These modules are integrated into the backbone network via a parallel canvas mechanism that injects them as additional attention tokens, eliminating the need for any architectural changes. This approach enables modular audio-visual control for the first time, allowing flexible composition of control signals and efficient training. On the VACE benchmark, the method outperforms existing approaches in depth- and pose-guided generation, image inpainting, and extrapolation tasks, achieves competitive performance in camera control and audio-visual synthesis, and significantly reduces training costs.

0 citationsRead paper

LTX-2: Efficient Joint Audio-Visual Foundation Model

Jan 06, 2026arXiv.org

This work addresses the prevailing limitation of existing text-to-video diffusion models in generating high-quality, synchronized audio that aligns semantically, emotionally, and atmospherically with the visual content. To this end, we propose a unified audio-visual generative foundation model featuring an asymmetric dual-stream Transformer architecture—comprising a 14B-parameter video stream and a 5B-parameter audio stream. The model leverages modality-aware classifier-free guidance (CFG), cross-modal AdaLN, and bidirectional audio-visual cross-attention mechanisms to achieve efficient co-generation and precise temporal alignment. Integrated with temporal positional encoding and a multilingual text encoder, our approach achieves state-of-the-art audio-visual quality and prompt fidelity within an open-source framework, matching the performance of closed-source counterparts while significantly reducing computational overhead and inference latency.

6 citationsRead paper

Enhancing Forecasting with a 2D Time Series Approach for Cohort-Based Data

Aug 21, 2025

This work addresses the challenge of cohort-based two-dimensional time series forecasting under few-shot learning conditions. Methodologically, it introduces a novel neural network framework that jointly models cohort-level structural patterns and temporal dynamics: (1) cohort embedding is innovatively integrated into 2D time series modeling to co-learn intra-cohort temporal evolution (across time) and inter-cohort heterogeneity (across cohorts); (2) a lightweight spatiotemporal attention mechanism is designed to enhance generalization under sparse data regimes. Extensive experiments on multi-source real-world datasets from finance and marketing domains demonstrate that the proposed model achieves average MAE improvements of 12.7%–23.4% over state-of-the-art baselines—including TCN, Transformer, and matrix factorization–based approaches—while significantly improving both predictive accuracy and business interpretability in low-data scenarios. The architecture exhibits strong practical deployability due to its parameter efficiency and robustness.

0 citationsRead paper