NeuroGuard: Neural Gradient Update Aware of Representation Damage

📅 2026-08-08
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenges of catastrophic forgetting of old classes and poor learning of new classes in long-tailed class-incremental learning (LT-CIL), which stem from severe data imbalance. To mitigate these issues without introducing additional parameters, the authors propose a task-boundary-aware gradient update modulation mechanism that jointly regulates feature representation updates through adaptive gradient scaling (AGS), confidence-ranked knowledge distillation reweighting (CRK), and a fragility-based entropy gate (FBE). This approach uniquely integrates teacher model uncertainty, prediction confidence ranking, and memory fragility to dynamically adjust both knowledge distillation weights and gradient magnitudes. Evaluated across five LT-CIL settings, the method consistently outperforms the DGR baseline and achieves state-of-the-art task-agnostic accuracy on four mainstream benchmarks, significantly improving performance across new, old, and medium-frequency classes.
📝 Abstract
Long-tailed class-incremental learning (LT-CIL) must learn new classes from imbalanced streams while retaining old classes. Existing methods mainly change replay, classifiers, or losses. We study a different factor, namely how strongly the feature representation should be updated at each task boundary. We propose NeuroGuard, an update-control method added to DGR, a replay-based LT-CIL baseline, without adding learnable parameters. NeuroGuard preserves DGR's replay memory, classifier, and set of loss terms. Adaptive Gradient Scaling (AGS) converts teacher uncertainty into one task-wise gradient scale. Confidence-Ranked Knowledge Distillation Reweighting (CRK) gives larger knowledge-distillation weights to replay samples that the teacher predicts less decisively. Fragility-Blended Entropy Gate (FBE) adds old-memory leakage to the scale decision. Across five LT-CIL settings, NeuroGuard improves over DGR in every setting. In the four main benchmark comparisons, it achieves the best task-agnostic accuracy among the compared methods. The gains extend to both old- and new-class accuracy, while medium-frequency accuracy improves consistently across all five settings. Controlled comparisons show that the gain does not come from generic gradient suppression: AGS outperforms a matched fixed-scale control in all five settings, demonstrating that boundary-specific scaling is more effective than applying the same average scale throughout learning.
Problem

Research questions and friction points this paper is trying to address.

long-tailed class-incremental learning
representation damage
task boundary
gradient update
class imbalance
Innovation

Methods, ideas, or system contributions that make the work stand out.

NeuroGuard
Adaptive Gradient Scaling
Confidence-Ranked Knowledge Distillation
Fragility-Blended Entropy Gate
Long-tailed Class-Incremental Learning
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
T
Taigo Sakai
Meijo University, 1-501 Shiogamaguchi, Tempaku-ku, Nagoya 468-8502, Japan
K
Kazuhito Hotta
Meijo University, 1-501 Shiogamaguchi, Tempaku-ku, Nagoya 468-8502, Japan