SubZero+: Efficient Zeroth-Order LLM Fine-Tuning via Large Learning Rates

📅 2026-08-16
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of high gradient estimation variance, unstable convergence, and learning rate sensitivity in zero-order fine-tuning of large language models by proposing the SubZero+ framework. This method innovatively integrates intra-layer low-rank subspace multi-query gradient estimation, an adaptive subspace Adam optimizer, and a QR sign correction mechanism to effectively eliminate directional ambiguity and expand the stable learning rate range. Experiments on models ranging from 1.3B to 32B parameters demonstrate that SubZero+ comprehensively outperforms existing zero-order baselines. It achieves performance comparable to first-order fine-tuning with minimal memory overhead, significantly enhancing both the stability and practicality of zero-order optimization for large-scale model adaptation.
📝 Abstract
Zeroth-order (ZO) optimization enables backpropagation-free fine-tuning of large language models, but existing ZO methods suffer from high-variance gradient estimators, making convergence unstable and highly sensitive to learning rates. We propose SubZero+, an improved SubZero framework that improves stability in three complementary ways: (i) multi-query gradient estimation within layer-specific low-rank subspaces to reduce variance without exhibiting the multi-query paradox; (ii) a subspace Adam optimizer that performs adaptive updates using in-subspace multi-query gradient statistics; and (iii) a sign correction for QR-based subspace construction to ensure Haar-distributed projection matrices, eliminating implementation-dependent orientation ambiguity. Experiments on models from 1.3B to 32B across SuperGLUE, under both full-parameter tuning and LoRA, show that SubZero+ consistently outperforms prior ZO baselines, enlarges the stable learning-rate range, and narrows the gap to first-order methods with minimal extra memory overhead.
Problem

Research questions and friction points this paper is trying to address.

Zeroth-Order Optimization
LLM Fine-Tuning
Gradient Estimation Variance
Convergence Stability
Learning Rate Sensitivity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Zeroth-Order Optimization
Low-Rank Subspace Gradient Estimation
Subspace Adam Optimizer
Haar-Distributed Projection
LLM Fine-Tuning
Z
Ziming Yu
Beijing Normal University
S
Shuyao Xiao
Beijing Normal University
Xingyu Zhao
Xingyu Zhao
Associate Professor, University of Warwick
Software ReliabilitySafe AIBayesian InferenceProbabilistic Model CheckingSafety Assurance
S
Sike Wang
Beijing Normal University
P
Pan Zhou
Singapore Management University
P
Peiyu Zang
Beijing Normal University
X
Xiangda Yan
Xiaomi Inc.
Y
Yongjie Yang
Xiaomi Inc.
J
Jia Li
Beijing Normal University; Beijing Key Laboratory of Artificial Intelligence for Education; Engineering Research Center of Intelligent Technology and Educational Application (MOE)