Preventing Model Collapse: A Fisher-Rao Perspective on the Dynamics of Training with Synthetic Data

📅 2026-09-16
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文通过Fisher-Rao度量方法分析了防止模型崩溃所需的最小真人数据比例,解决了大语言模型在合成数据训练中出现的退化问题。
📝 Abstract
Large Language Models (LLMs) are now routinely trained using synthetic data, since high-quality human data has been exhausted by the ever increasing needs of larger and larger models. However, recursive training on synthetic data frequently induces model collapse, a degenerative feedback loop where models progressively forget the true underlying data distribution. Training on a mixture of synthetic and fresh human data is a logical countermeasure and can prevent model collapse. However, it is an open question as to what is the exact minimum required ratio of human-to-synthetic data to maintain training stability. In this paper, we establish rigorous theoretical guarantees on the minimum rate of human data required to prevent model collapse. Although previous work established a formal lower bound for this ratio, such bound can be vacuous for very high dimensions, as the analysis relies on the usual Euclidean metric in R^n and is not adapted to the space of categorical probability distributions. Instead, in this paper we explicitly leverage the information-geometric structure of the probability simplex by analyzing the dynamics of the process under the Fisher-Rao metric. We derive quantitative contraction and invariance bounds that are stable and do not become trivial as the dimensions increase. Thus, we show that the effective required data ratio to prevent model collapse is different than previously implied.
Problem

Research questions and friction points this paper is trying to address.

model collapse
synthetic data
human data
training stability
Fisher-Rao metric
Innovation

Methods, ideas, or system contributions that make the work stand out.

Fisher-Rao metric
information-geometric structure
model collapse
synthetic data
human data ratio
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
M
Matteo Marchi
Electrical and Computer Engineering Department, University of California at Los Angeles, Los Angeles, CA 90095 USA
J
João Pedro Silvestre
Electrical and Computer Engineering Department, University of California at Los Angeles, Los Angeles, CA 90095 USA
Bahman Gharesifard
Bahman Gharesifard
Professor of Mathematics at Queen's University
Control TheoryOptimizationReinforcement LearningNeural NetworksGeometric Control
P
Paulo Tabuada
Electrical and Computer Engineering Department, University of California at Los Angeles, Los Angeles, CA 90095 USA