Institution profile

Luma AI

Industry researchnorthamerica · us
Official website
Research library16linked papers
Opportunities0open roles
Selected work

Representative Papers

Inductive Moment Matching

Mar 10, 2025

Diffusion and flow-matching models achieve high-quality generation but suffer from slow inference; distillation to few-step sampling often leads to instability and requires extensive hyperparameter tuning. To address this, we propose Inductive Moment Matching (IMM), a novel generative paradigm that enables one- or few-step sampling without pretraining and with single-stage end-to-end training. IMM introduces the first moment-matching mechanism with distribution-level convergence guarantees—eliminating the need for teacher models or multi-stage optimization. It relies solely on a principled, moment-based objective function and standard neural architectures. On ImageNet-256, IMM achieves a FID of 1.99 using only eight sampling steps; on CIFAR-10, it attains a FID of 1.98 in just two steps—setting a new state-of-the-art for zero-initialized few-step generative modeling.

1 citationsRead paper

Inference-Time Scaling for Joint Audio-Video Generation

Jun 02, 2026

This work addresses the challenges in joint audio-visual generation, where semantic alignment, perceptual quality, and audio-visual synchronization are difficult to optimize simultaneously, often requiring prohibitively expensive training. The study introduces inference-time scaling to multimodal generation for the first time, proposing a training-free multi-objective optimization framework. By leveraging collaborative guidance from multiple verifiers and an adaptive reward weighting (ARW) mechanism, the method dynamically aggregates online reward signals during inference to achieve balanced improvements across generation objectives. Experiments on VGGSound and JavisBench-mini demonstrate that the approach significantly enhances semantic consistency, audio-visual synchronization, and overall perceptual quality of generated content without any additional training.

0 citationsRead paper

VisAnalog: A Diagnostic Suite for Visual Concept Transfer on Natural Images

May 21, 2026

This work investigates whether visual models can transfer and manipulate concept-level attributes in natural images beyond merely recognizing static concepts. To this end, we introduce VisAnalog, the first multi-step visual analogy benchmark for natural images, comprising 617 human-verified analogy questions (A:B::C:?) constructed via programmatically controlled transformations such as scaling, flipping, and hue rotation. We further propose a program-conditioned evaluation protocol to disentangle errors arising from relational reasoning versus transformation execution. Experimental results reveal that state-of-the-art vision-language models achieve substantially lower accuracy than humans, with performance degrading sharply as the number of transformation steps increases, highlighting relational reasoning as a key bottleneck.

0 citationsRead paper

GeoNVS: Geometry Grounded Video Diffusion for Novel View Synthesis

Mar 16, 2026

Existing video diffusion models often suffer from geometric distortions and limited camera controllability in novel view synthesis. To address this, this work proposes GS-Adapter, a plug-and-play framework that explicitly integrates 3D Gaussian Splatting geometry priors with diffusion features in the latent space to enforce geometric consistency during view generation. The method requires no modifications to the input pipeline or additional training and is compatible with diverse geometric representations. Evaluated across nine scenes and eighteen configurations, GS-Adapter achieves state-of-the-art performance, outperforming SEVA and CameraCtrl by 11.3% and 14.9%, respectively. It reduces translational error by up to a factor of two and decreases Chamfer distance by as much as sevenfold, demonstrating significantly improved geometric fidelity and view coherence.

0 citationsRead paper
Recent publications

Latest Papers

Inference-Time Scaling for Joint Audio-Video Generation

Jun 02, 2026

This work addresses the challenges in joint audio-visual generation, where semantic alignment, perceptual quality, and audio-visual synchronization are difficult to optimize simultaneously, often requiring prohibitively expensive training. The study introduces inference-time scaling to multimodal generation for the first time, proposing a training-free multi-objective optimization framework. By leveraging collaborative guidance from multiple verifiers and an adaptive reward weighting (ARW) mechanism, the method dynamically aggregates online reward signals during inference to achieve balanced improvements across generation objectives. Experiments on VGGSound and JavisBench-mini demonstrate that the approach significantly enhances semantic consistency, audio-visual synchronization, and overall perceptual quality of generated content without any additional training.

0 citationsRead paper

VisAnalog: A Diagnostic Suite for Visual Concept Transfer on Natural Images

May 21, 2026

This work investigates whether visual models can transfer and manipulate concept-level attributes in natural images beyond merely recognizing static concepts. To this end, we introduce VisAnalog, the first multi-step visual analogy benchmark for natural images, comprising 617 human-verified analogy questions (A:B::C:?) constructed via programmatically controlled transformations such as scaling, flipping, and hue rotation. We further propose a program-conditioned evaluation protocol to disentangle errors arising from relational reasoning versus transformation execution. Experimental results reveal that state-of-the-art vision-language models achieve substantially lower accuracy than humans, with performance degrading sharply as the number of transformation steps increases, highlighting relational reasoning as a key bottleneck.

0 citationsRead paper

GeoNVS: Geometry Grounded Video Diffusion for Novel View Synthesis

Mar 16, 2026

Existing video diffusion models often suffer from geometric distortions and limited camera controllability in novel view synthesis. To address this, this work proposes GS-Adapter, a plug-and-play framework that explicitly integrates 3D Gaussian Splatting geometry priors with diffusion features in the latent space to enforce geometric consistency during view generation. The method requires no modifications to the input pipeline or additional training and is compatible with diverse geometric representations. Evaluated across nine scenes and eighteen configurations, GS-Adapter achieves state-of-the-art performance, outperforming SEVA and CameraCtrl by 11.3% and 14.9%, respectively. It reduces translational error by up to a factor of two and decreases Chamfer distance by as much as sevenfold, demonstrating significantly improved geometric fidelity and view coherence.

0 citationsRead paper

Terminal Velocity Matching

Nov 24, 2025

To address the challenges in high-fidelity one-step/few-step generation—namely, the difficulty of flow matching in modeling transitions across diffusion time steps and the absence of terminal behavioral constraints—this paper proposes Terminal Velocity Matching (TVM). TVM is the first to shift flow matching regularization from the initial to the terminal time, with theoretical proof that it upper-bounds the 2-Wasserstein distance. Methodologically, we design a lightweight Diffusion Transformer architecture incorporating a fused attention kernel to efficiently support Jacobian-vector product backpropagation, and adopt a single-stage training strategy. On ImageNet-256×256, our method achieves FID scores of 3.29 (one-step) and 1.99 (four-step); on ImageNet-512×512, it attains 4.32 (one-step) and 2.94 (four-step), surpassing current state-of-the-art methods.

0 citationsRead paper