Hypernym Bias: Unraveling Deep Classifier Training Dynamics through the Lens of Class Hierarchy
This work investigates how deep classifiers dynamically model semantic hierarchies among classes during training. To address this, we propose the first training dynamics analysis framework grounded in class hierarchy evolution. Our method integrates hierarchy-aware visualization, cross-layer representation tracking, label clustering metrics, and quantification of neural collapse. It reveals that feature manifolds progressively align with hypernym–hyponym semantic structures across network layers: early layers prioritize separation of coarse-grained (hypernym) categories, while deeper layers refine fine-grained (hyponym) distinctions; moreover, neural collapse occurs earlier in the hypernym label space. Experiments demonstrate that deep networks inherently perform hierarchical learning aligned with the intrinsic data hierarchy, yielding feature representations highly consistent with ground-truth class taxonomy. This significantly enhances the interpretability of training dynamics on standard benchmarks such as ImageNet.