🤖 AI Summary
This study addresses the lack of systematic characterization in deep network training trajectories by proposing short-horizon predictability as a retrospective diagnostic metric. Utilizing diverse probes, including displacement direction and subspace residuals, alongside a zero-calibration readout mechanism, this work quantifies the temporal redundancy and structural features of parameter updates. The research reveals distinct dynamical differences between auxiliary and primary parameters, elucidates the spatiotemporal distribution of predictability, and validates the modulatory effects of architecture and training strategies on trajectory structure. Collectively, these findings provide fine-grained quantitative tools and novel perspectives for understanding neural network training dynamics, offering a robust framework for analyzing the intrinsic geometric and temporal properties of optimization paths in deep learning systems.
📝 Abstract
Modern deep networks are trained through long update trajectories, yet their temporal organization remains less systematically characterized than architectures, losses, or optimizers. We study short-horizon predictability as a measure of temporal redundancy: where, when, and under which training conditions recent updates contain information about near-future parameter motion. We combine three complementary probe families, displacement-direction, subspace-residual, and predictor-based probes, with convention-aware, null-calibrated group-level readouts, and apply them to multi-pass vision training on CIFAR and public Pythia pretraining checkpoints. Across both regimes, vector-like tensors such as normalization parameters and biases (auxiliary parameters) exhibit simpler short-horizon dynamics than matrix-like feature-transforming weights (bulk parameters), whose predictable behavior concentrates in localized, time-varying pockets. Agreement within and across probe families, and with independent trajectory diagnostics, indicates that these measurements capture intrinsic trajectory structure, while probe differences distinguish complementary forms of temporal organization. Controlled CIFAR comparisons further show that architecture and training recipe systematically modulate the measured structure. A Pythia-70M case study further exposes a sequence of role-, depth-, and scale-dependent events, including bulk ESA falling below the random sign-agreement level and the emergence and redistribution of predictable qkv pockets across layers. These results position short-horizon predictability as a retrospective, parameter-resolved diagnostic of training dynamics.