Fenchel-Young Duality Gaps: Certified Early Stopping for Regularized Inverse Problems

📅 2026-09-15
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文研究了正则化逆问题中的可计算误差界和认证的提前停止方法,通过Fenchel-Young对偶间隙将总误差分解为数据保真度损失和正则化损失,提供了一种基于对偶间隙下降的提前停止规则。
📝 Abstract
We study computable error bounds and certified early stopping for regularized inverse problems, where a data-fidelity term is traded against a regularizer. The analysis relies on an exact duality-gap identity that splits the total gap of $F(Φμ)+λR(μ)$ into a data-fidelity Fenchel--Young loss and a regularizer Fenchel--Young loss, $ Δ(μ,h)=L_F(Φμ\parallel h)+λL_R(μ\parallelη),\qquad η=-Φ^\star h/λ, $ valid for any primal point $μ$ and any dual point $h$. The data-fidelity term $F$ is strictly convex, so wherever $F^\star$ is differentiable the loss $L_F(Φμ\parallel h)$ is the Bregman divergence of~$F$ between the prediction $Φμ$ and $\nabla F^\star(h)$, and it vanishes exactly at \emph{Mirror Alignment} $h=\nabla F(Φμ)$. Evaluated at a dual-feasible point $\tilde h$, the gap~$Δ(μ,\tilde h)$ is computable and \emph{oracle-free}, meaning that it uses no knowledge of the solution, and it bounds the suboptimality of $μ$. Under the source condition, the same Fenchel--Young losses give \emph{a priori} bounds on the estimation and prediction errors. Their scale is the irreducible model and noise error $L_F(Φμ^\star\parallel h^\star)$, which vanishes exactly when Mirror Alignment holds at the certificate. This gives an early-stopping rule: run the algorithm until the regularizer Fenchel--Young loss falls below a tolerance $ε$. A constructive version of the Brøndsted--Rockafellar theorem then turns the current pair into an exact \emph{dual-feasible} one, and this proxy lifts to an exact primal certificate. We build the proxy by a proximal step in the geometry of the fidelity, with Bregman kernel $F^\star$ and tilted by the prediction $Φμ$: it recovers the Euclidean step of Carlier when $F$ is the squared error, and it reduces the duality gap by the regularizer Fenchel--Young loss, up to a second-order remainder that vanishes in the quadratic case. Our running example is the Generalized Beurling--Lasso (GBL), where $R$ is the total-variation norm on signed measures. It contains the classical Beurling--Lasso, obtained with the squared error, and also covers robust, logistic, entropic and inverse-optimal-transport losses. The same duality gap certifies deep-learning optimizers such as Lion-K and Muon, in their proximal form, as solvers of the regularized program. A companion paper by the same authors builds on these error bounds to establish exact support recovery for the GBL under a non-degenerate source condition.
Problem

Research questions and friction points this paper is trying to address.

regularized inverse problems
duality gap
early stopping
Fenchel-Young loss
data-fidelity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Fenchel-Young Duality
Certified Early Stopping
Regularized Inverse Problems
Bregman Divergence
Mirror Alignment