Compute-Optimal Pretrain--Fine-tune in Ridge Gradient Descent

📅 2026-09-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文研究了在固定训练预算下,如何最优分配预训练和微调的计算资源问题,通过正则化最小二乘法和梯度下降方法来解决。
📝 Abstract
Pretraining followed by fine-tuning introduces a compute-allocation problem: under a fixed training budget, compute spent improving the upstream objective reduces the compute available for downstream adaptation. Despite its practical importance, this trade-off is not yet well understood theoretically, even in simple models. In this paper, we cast this allocation as a compute-split problem under a two-stage pretrain--fine-tune procedure with fixed total optimisation budget, using regularised least squares trained by gradient descent as a tractable setting. We characterise the optimal split under data-dependent evaluation geometries induced by the fine-tuning problem. Our results show that the allocation depends on how pretraining directions affect fine-tuning predictions and how fine-tuning shifts are seen through downstream data geometry. In particular, the relevant quantities are determined by prediction-relevant spectral components of the pretraining and fine-tuning empirical covariances. Technically, the analysis relies on a basis-invariant, eigenspace-level spectral decomposition, together with perturbative control of the non-commuting pretraining and fine-tuning dynamics.
Problem

Research questions and friction points this paper is trying to address.

pretrain-fine-tune
compute-allocation
optimisation budget
regularised least squares
gradient descent
Innovation

Methods, ideas, or system contributions that make the work stand out.

compute-allocation
pretrain--fine-tune
spectral decomposition
gradient descent
data geometry
🔎 Similar Papers
No similar papers found.
A
Alex Buna
Department of Statistics, University of Oxford, UK
Fanghui Liu
Fanghui Liu
Assistant Professor, University of Warwick
Foundations of Modern ML
P
Patrick Rebeschini
Department of Statistics, University of Oxford, UK