Active-Trace Complexity Bounds for Moreau--Yosida Unadjusted Langevin Sampling

📅 2026-08-13
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the inefficiency of the traditional MYULA algorithm for sampling from nonsmooth composite target distributions, which stems from its reliance on a global curvature bound. To overcome this limitation, the authors propose an error-control mechanism based on a reference active trace, replacing the global curvature bound to more accurately capture the dominant discretization error. By integrating techniques including Moreau envelope smoothing, weak Hessian trace analysis, heat semi-step trajectory averaging, and curvature-tube estimation, the method achieves—for the first time—an $\widetilde{O}(\varepsilon^{-2})$ iteration complexity under structured nonsmooth penalties such as Lasso, group, and total variation regularization, improving upon the generic $\widetilde{O}(\varepsilon^{-3})$ bound. The approach also provides end-to-end Wasserstein error guarantees, attaining sampling efficiency comparable to that in the smooth setting.
📝 Abstract
We study the Moreau--Yosida unadjusted Langevin algorithm (MYULA) for the nonsmooth composite target \[ π(dx)\propto \exp\{-f(x)-g(x)\}\,dx, \qquad x\in\mathbb R^d, \] where \(f\) is \(m\)-strongly convex with \(L_f\)-Lipschitz gradient and \(g\) is convex and \(G\)-Lipschitz. Let \(g_λ\) be the Moreau envelope of \(g\), \(π_λ\) the corresponding smoothed target, and \(a_λ=\operatorname{tr}H_λ\), where \(H_λ\) is the a.e./weak Hessian of \(g_λ\). We show that the leading MYULA discretization error is controlled by the reference active trace \(B_{\mathrm{ref}}\), the average of \(a_λ\) along the heat substep of one MYULA update started from \(π_λ\), rather than by the global curvature bound \(d/λ\). If \(M_λ\) is an a.e. upper bound for \(a_λ\), then, up to logarithmic factors, \[ N \lesssim \frac{1}{m} \left[ L_f + \frac{ τ_f+G^2+B_{\mathrm{ref}} }{ \varepsilon_{\mathrm{alg}}^2 } + \frac{M_λ}{\varepsilon_{\mathrm{alg}}} \right], \qquad τ_f:= \sup_x\operatorname{tr}\nabla^2 f(x), \] iterations suffice to ensure \(\sqrt m\,W_2(μ_N,π_λ)\leq\varepsilon_{\mathrm{alg}}\), where \(μ_N\) is the law of the \(N\)-th iterate and \(W_2\) is the quadratic Wasserstein distance. We also prove the Moreau-bias bound \[ \sqrt m\,W_2(π_λ,π) \leq \frac{G^2λ}{4}. \] Thus, choosing \(λ\asymp\varepsilon/G^2\) gives an end-to-end guarantee for \(π\). The universal estimate \(B_{\mathrm{ref}}\leq d/λ\) yields \(\widetilde O(\varepsilon^{-3})\) accuracy dependence. For the structured piecewise-linear, lasso-type, group, and total-variation penalties considered here, curvature--tube estimates make \(B_{\mathrm{ref}}\) independent of \(λ\), yielding \(\widetilde O(\varepsilon^{-2})\) for the same classical MYULA kernel.
Problem

Research questions and friction points this paper is trying to address.

nonsmooth sampling
Moreau--Yosida regularization
Langevin algorithm
discretization error
composite target distribution
Innovation

Methods, ideas, or system contributions that make the work stand out.

Moreau-Yosida regularization
active trace
Langevin sampling
nonsmooth composite sampling
Wasserstein convergence
💼 Related Jobs
No related jobs found.