A Theory of Initialisation's Impact on Specialisation
This work challenges the necessity of neuron specialization for mitigating catastrophic forgetting in continual learning, revealing that specialization is primarily governed by network initialization rather than intrinsic task properties. Method: Through theoretical analysis and empirical validation, the authors demonstrate that weight imbalance and high weight entropy actively induce localized representations; they provide the first theoretical proof that specialization is not inherent but contingent. They further derive a quantitative relationship between specialization degree and initialization parameters, and reproduce the monotonic relationship between task similarity and forgetting rate even in non-specialized networks. Contribution/Results: Specialized initialization significantly enhances Elastic Weight Consolidation (EWC) performance, an effect attributable to initialization-induced prior shaping of representation structure. These findings establish a novel theoretical foundation for regularization design in continual learning and yield principled guidelines for initialization strategy selection.