A Dynamic Knowledge Distillation Method Based on the Gompertz Curve

📅 2025-10-24
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Traditional knowledge distillation overlooks the dynamic evolution of student models’ cognitive capacity, limiting knowledge transfer efficiency. To address this, we propose Gompertz-CNN—a novel distillation framework that explicitly models the sigmoidal (S-shaped) learning dynamics of students by integrating the Gompertz growth model into the distillation process for the first time. We further design a phase-aware, time-varying loss weighting mechanism that jointly optimizes Wasserstein-based feature alignment and gradient propagation matching, enabling coordinated alignment at both the feature and backward-propagation levels. The resulting framework is end-to-end trainable and supports multi-objective dynamic distillation. Extensive experiments on CIFAR-10 and CIFAR-100 demonstrate significant improvements over state-of-the-art distillation methods, achieving up to 8% and 4% absolute accuracy gains, respectively. Moreover, Gompertz-CNN exhibits strong effectiveness and robustness across diverse teacher–student architecture pairs.

Technology Category

Application Category

📝 Abstract
This paper introduces a novel dynamic knowledge distillation framework, Gompertz-CNN, which integrates the Gompertz growth model into the training process to address the limitations of traditional knowledge distillation. Conventional methods often fail to capture the evolving cognitive capacity of student models, leading to suboptimal knowledge transfer. To overcome this, we propose a stage-aware distillation strategy that dynamically adjusts the weight of distillation loss based on the Gompertz curve, reflecting the student's learning progression: slow initial growth, rapid mid-phase improvement, and late-stage saturation. Our framework incorporates Wasserstein distance to measure feature-level discrepancies and gradient matching to align backward propagation behaviors between teacher and student models. These components are unified under a multi-loss objective, where the Gompertz curve modulates the influence of distillation losses over time. Extensive experiments on CIFAR-10 and CIFAR-100 using various teacher-student architectures (e.g., ResNet50 and MobileNet_v2) demonstrate that Gompertz-CNN consistently outperforms traditional distillation methods, achieving up to 8% and 4% accuracy gains on CIFAR-10 and CIFAR-100, respectively.
Problem

Research questions and friction points this paper is trying to address.

Dynamic knowledge distillation adapting to student learning stages
Overcoming limitations of static knowledge transfer methods
Improving accuracy through stage-aware distillation strategy
Innovation

Methods, ideas, or system contributions that make the work stand out.

Dynamic distillation loss weighting via Gompertz curve
Feature alignment using Wasserstein distance measurement
Gradient matching for backward propagation alignment
🔎 Similar Papers
No similar papers found.
H
Han Yang
SmartCity College, Beijing Union University, Beijing 100101, China
G
Guangjun Qin
SmartCity College, Beijing Union University, Beijing 100101, China