Decomposing the Depth Profile of Fine-Tuning

📅 2026-04-18
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study investigates whether the depth-wise distribution of representational changes during fine-tuning arises from intrinsic model properties or the magnitude of gradient flow. Through 240 fine-tuning experiments spanning serial and parallel architectures and model scales from 125M to 6.9B parameters, combined with representational similarity analysis, layer-wise relative weight change (|ΔW|/|W|), and task-agnostic target distance metrics, the work systematically reveals that fine-tuning depth profiles are jointly governed by architecture type, task objective, and model scale. Under standard training, representational shifts concentrate in upper layers; under controlled conditions, only serial architectures retain a positive slope in small models, while parallel architectures maintain it solely in causal language modeling. Above 1.3B parameters, both architectures exhibit positive slopes, with profile steepness correlating with initial target distance and width primarily dictated by architecture—demonstrating that “localized gradients” emerge from a multifactorial interplay.

Technology Category

Application Category

📝 Abstract
Fine-tuning adapts pretrained networks to new objectives. Whether the resulting depth profile of representational change reflects an intrinsic property of the model or the magnitude of gradient flow has not been tested directly. We measure this profile across 240 fine-tuning runs spanning 15 models in four architecture families (encoder and decoder transformers, a state-space model, and an RNN) at scales from 125M to 6.9B parameters. Representational change concentrates in output-proximal layers in every standard-training run except one. We apply a per-layer control that equalizes $\|ΔW\|/\|W\|$ across layers after each optimizer step. Under this control, the profile persists in some conditions and collapses in others. At 125M--350M, sequential-block architectures (BERT, OPT, GPT-2) retain the slope across tested objectives while parallel-block architectures (Pythia, CodeGen) retain it only for causal-language-modeling objectives. This architectural distinction narrows at 1.3B--1.4B, where both block types show positive equal-step slopes for CausalLM. Under standard training, profile shape is described by two additional axes: steepness tracks a training-free objective distance at initialization, and profile width is dominated by architecture. We treat the locality gradient, the depthwise slope of representational change, as a composite phenomenon whose components are scale-dependent.
Problem

Research questions and friction points this paper is trying to address.

fine-tuning
depth profile
representational change
gradient flow
architecture
Innovation

Methods, ideas, or system contributions that make the work stand out.

depth profile
fine-tuning dynamics
locality gradient
architecture dependence
scale-dependent representation
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
J
Jayadev Billa
Unaffiliated researcher; previously at ISI@USC, Yahoo, Nuance, and BBN