Hypernym Bias: Unraveling Deep Classifier Training Dynamics through the Lens of Class Hierarchy

๐Ÿ“… 2025-02-17
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This work investigates how deep classifiers dynamically model semantic hierarchies among classes during training. To address this, we propose the first training dynamics analysis framework grounded in class hierarchy evolution. Our method integrates hierarchy-aware visualization, cross-layer representation tracking, label clustering metrics, and quantification of neural collapse. It reveals that feature manifolds progressively align with hypernymโ€“hyponym semantic structures across network layers: early layers prioritize separation of coarse-grained (hypernym) categories, while deeper layers refine fine-grained (hyponym) distinctions; moreover, neural collapse occurs earlier in the hypernym label space. Experiments demonstrate that deep networks inherently perform hierarchical learning aligned with the intrinsic data hierarchy, yielding feature representations highly consistent with ground-truth class taxonomy. This significantly enhances the interpretability of training dynamics on standard benchmarks such as ImageNet.

Technology Category

Application Category

๐Ÿ“ Abstract
We investigate the training dynamics of deep classifiers by examining how hierarchical relationships between classes evolve during training. Through extensive experiments, we argue that the learning process in classification problems can be understood through the lens of label clustering. Specifically, we observe that networks tend to distinguish higher-level (hypernym) categories in the early stages of training, and learn more specific (hyponym) categories later. We introduce a novel framework to track the evolution of the feature manifold during training, revealing how the hierarchy of class relations emerges and refines across the network layers. Our analysis demonstrates that the learned representations closely align with the semantic structure of the dataset, providing a quantitative description of the clustering process. Notably, we show that in the hypernym label space, certain properties of neural collapse appear earlier than in the hyponym label space, helping to bridge the gap between the initial and terminal phases of learning. We believe our findings offer new insights into the mechanisms driving hierarchical learning in deep networks, paving the way for future advancements in understanding deep learning dynamics.
Problem

Research questions and friction points this paper is trying to address.

Analyzing deep classifier training dynamics
Tracking evolution of class hierarchy
Bridging initial and terminal learning phases
Innovation

Methods, ideas, or system contributions that make the work stand out.

Hierarchical class evolution tracking
Feature manifold evolution framework
Neural collapse properties analysis
๐Ÿ”Ž Similar Papers
R
Roman Malashin
Saint-Petersburg State University of Aerospace Instrumentation
V
Valeria Yachnaya
Pavlov Institute of Physiology, RAS
A
Alexander Mullin
Saint-Petersburg State University of Aerospace Instrumentation