A Geometrically-Grounded Drive for MDL-Based Optimization in Deep Learning

📅 2026-03-12
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work proposes a novel approach that actively embeds the Minimum Description Length (MDL) principle into the optimization process of deep learning to jointly enhance model simplicity and generalization. By constructing a cognitive manifold driven by coupled Ricci flows, the method introduces a geometry-driven MDL Drive mechanism that dynamically compresses internal representations during training, balancing fidelity and complexity. Theoretical analysis establishes that this mechanism guarantees a monotonically decreasing description length, undergoes a finite number of topological phase transitions, and exhibits universal critical behavior. With a per-iteration computational complexity of O(N log N), the algorithm automatically simplifies model architecture, improves generalization, and demonstrates numerical stability alongside exponential convergence in empirical evaluations.

Technology Category

Application Category

📝 Abstract
This paper introduces a novel optimization framework that fundamentally integrates the Minimum Description Length (MDL) principle into the training dynamics of deep neural networks. Moving beyond its conventional role as a model selection criterion, we reformulate MDL as an active, adaptive driving force within the optimization process itself. The core of our method is a geometrically-grounded cognitive manifold whose evolution is governed by a \textit{coupled Ricci flow}, enriched with a novel \textit{MDL Drive} term derived from first principles. This drive, modulated by the task-loss gradient, creates a seamless harmony between data fidelity and model simplification, actively compressing the internal representation during training. We establish a comprehensive theoretical foundation, proving key properties including the monotonic decrease of description length (Theorem~\ref{thm:convergence}), a finite number of topological phase transitions via a geometric surgery protocol (Theorems~\ref{thm:surgery}, \ref{thm:ultimate_fate}), and the emergence of universal critical behavior (Theorem~\ref{thm:universality}). Furthermore, we provide a practical, computationally efficient algorithm with $O(N \log N)$ per-iteration complexity (Theorem~\ref{thm:complexity}), alongside guarantees for numerical stability (Theorem~\ref{thm:stability}) and exponential convergence under convexity assumptions (Theorem~\ref{thm:convergence_rate}). Empirical validation on synthetic regression and classification tasks confirms the theoretical predictions, demonstrating the algorithm's efficacy in achieving robust generalization and autonomous model simplification. This work provides a principled path toward more autonomous, generalizable, and interpretable AI systems by unifying geometric deep learning with information-theoretic principles.
Problem

Research questions and friction points this paper is trying to address.

Minimum Description Length
Deep Learning Optimization
Model Simplification
Generalization
Geometric Deep Learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Minimum Description Length (MDL)
Geometric Deep Learning
Ricci Flow
Model Simplification
Optimization Dynamics
M
Ming Lei
School of Aeronautics and Astronautics, Shanghai JiaoTong University, Shanghai China
S
Shufan Wu
School of Aeronautics and Astronautics, Shanghai JiaoTong University, Shanghai China
Christophe Baehr
Christophe Baehr
MÊtÊo-France