Carré du champ flow matching: better quality-generalisation tradeoff in generative models

📅 2025-10-07

📈 Citations: 0

✨ Influential: 0

career value

191K/year

🤖 AI Summary

Deep generative models face a fundamental trade-off between sample fidelity and generalization: high-fidelity generation often leads to memorization of training data rather than learning the underlying manifold structure. To address this, we propose geometrically aware noise regularization within the flow matching framework, replacing conventional homogeneous isotropic Gaussian noise with locally manifold-adapted anisotropic Gaussian noise. We theoretically establish its optimality and scalability. By modeling spatially varying covariance, our method is architecture-agnostic—compatible with MLPs, CNNs, and Transformers—and inherently suited to low-data and non-uniformly sampled regimes. Extensive experiments across diverse domains—including synthetic manifolds, point clouds, single-cell genomics, motion capture, and natural images—demonstrate consistent and significant improvements over standard flow matching. Notably, gains are most pronounced under data scarcity and non-uniform sampling, where generalization performance is markedly enhanced.

Technology Category

Application Category

📝 Abstract

Deep generative models often face a fundamental tradeoff: high sample quality can come at the cost of memorisation, where the model reproduces training data rather than generalising across the underlying data geometry. We introduce Carré du champ flow matching (CDC-FM), a generalisation of flow matching (FM), that improves the quality-generalisation tradeoff by regularising the probability path with a geometry-aware noise. Our method replaces the homogeneous, isotropic noise in FM with a spatially varying, anisotropic Gaussian noise whose covariance captures the local geometry of the latent data manifold. We prove that this geometric noise can be optimally estimated from the data and is scalable to large data. Further, we provide an extensive experimental evaluation on diverse datasets (synthetic manifolds, point clouds, single-cell genomics, animal motion capture, and images) as well as various neural network architectures (MLPs, CNNs, and transformers). We demonstrate that CDC-FM consistently offers a better quality-generalisation tradeoff. We observe significant improvements over standard FM in data-scarce regimes and in highly non-uniformly sampled datasets, which are often encountered in AI for science applications. Our work provides a mathematical framework for studying the interplay between data geometry, generalisation and memorisation in generative models, as well as a robust and scalable algorithm that can be readily integrated into existing flow matching pipelines.

Problem

Research questions and friction points this paper is trying to address.

Improves generative model tradeoff between quality and generalization

Replaces isotropic noise with geometry-aware anisotropic noise

Enhances performance on non-uniform and data-scarce scientific datasets

Innovation

Methods, ideas, or system contributions that make the work stand out.

Introduces geometry-aware anisotropic noise regularization

Replaces isotropic noise with spatially varying covariance

Optimally estimates scalable geometric noise from data

🔎 Similar Papers

No similar papers found.