Stability and Generalization of Straight-Through Estimators for Training Two-Layer Quantized Neural Networks

📅 2026-09-06
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文通过统计学习理论研究了二值激活两层神经网络中直通估计器的稳定性和泛化能力,利用算法稳定性解释其统计泛化,并给出了明确的模型稳定性和泛化界。
📝 Abstract
We study the identity straight-through estimator (STE) for training a two-layer binary-activation network with hinge loss from the perspective of Statistical Learning Theory (SLT). Our central question is whether algorithmic stability can explain the statistical generalization of the estimator produced by the discontinuous STE training rule. In the saturated-output regime, the zero-initialized samplewise STE recursion is exactly the stochastic subgradient descent on the convex latent loss $(-yu^\top x)_+$. This representation makes a stability analysis possible. We derive an exact distance identity for two coupled updates and prove approximate non-expansiveness of the common-example map, with a quadratic defect only when the two latent margins straddle zero. We then obtain explicit $\ell_2$ on-average model-stability and generalization bounds, transferring stability isometrically from the latent vector to the full first-layer matrix. Combining stability with a standard optimization bound yields an explicit excess induced-risk guarantee and the rate $O(n^{-1/2})$ when $T=n^2$. Under margin separability, a complementary argument gives the optimal-order $O(R^2/(\gamma^2n))$ expected excess misclassification error for a randomized one-pass STE iterate and a corresponding majority-vote bound.
Problem

Research questions and friction points this paper is trying to address.

algorithmic stability
statistical generalization
straight-through estimator
binary-activation network
hinge loss
Innovation

Methods, ideas, or system contributions that make the work stand out.

straight-through estimator
statistical learning theory
algorithmic stability
latent loss
generalization bounds
🔎 Similar Papers
No similar papers found.