UniDFKD: A Unified Semantic Prior Framework for Architecture-Agnostic Data-Free Knowledge Distillation

📅 2026-08-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitations of existing data-free knowledge distillation methods, which rely heavily on architecture-specific statistical priors such as batch normalization—leading to poor semantic quality of synthesized data and significant performance degradation when applied to modern architectures like Vision Transformers (ViTs) that lack such priors. To overcome this, the paper proposes UniDFKD, a unified, architecture-agnostic framework that introduces explicit semantic priors through three key mechanisms: language embedding–driven class semantic modulation, Gaussian prior–guided spatial semantic anchoring, and joint alignment of teacher–student feature-space evidence and predictions. This approach enables high-fidelity data synthesis and effective knowledge transfer across diverse architectures. Extensive experiments demonstrate substantial improvements over prior methods, with UniDFKD achieving an average absolute accuracy gain of over 20% in both homogeneous and heterogeneous settings, establishing a new state of the art in data-free knowledge distillation.
📝 Abstract
Data-Free Knowledge Distillation (DFKD) transfers knowledge from a pretrained teacher model to a compact student model by synthesizing semantically informative data, eliminating the need for access to the original training dataset. Existing DFKD methods rely heavily on architecture-specific statistical priors (e.g., Batch Normalization statistics) to guide data synthesis, however, such architecture-dependent priors are often absent in modern architectures such as Vision Transformers (ViTs), resulting in degraded semantic quality of the synthesized data and consequently catastrophic performance degradation. In this paper, we propose \emph{UniDFKD}, a unified data-free knowledge distillation framework that replaces architecture-specific statistics with explicit, architecture-agnostic semantic priors. \emph{UniDFKD} governs the entire synthesis-distillation pipeline along three dimensions: (1) Categorical Semantic Conditioning (CSC) defines \emph{what} to synthesize by persistently modulating the generator with language-derived embeddings to capture semantic diversity; (2) Spatial Semantic Anchoring (SSA) dictates \emph{where} evidence belongs by anchoring the teacher's spatial attributions to a Gaussian prior; and (3) Spatial Semantic Distillation (SSD) controls \emph{how} knowledge is transferred by explicitly aligning teacher-student spatial evidence alongside predictions. Extensive experiments across CNNs and ViTs demonstrate that UniDFKD establishes a new state-of-the-art, outperforming existing methods by an average absolute margin of over 20\% in both homogeneous and heterogeneous settings.
Problem

Research questions and friction points this paper is trying to address.

Data-Free Knowledge Distillation
Architecture-Agnostic
Semantic Priors
Vision Transformers
Knowledge Distillation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Data-Free Knowledge Distillation
Architecture-Agnostic
Semantic Prior
Spatial Attribution
Vision Transformers
🔎 Similar Papers
No similar papers found.
X
Xuewan He
University of Electronic Science and Technology of China, Chengdu, China
T
Tong Chu
University of Electronic Science and Technology of China, Chengdu, China
Z
Zihan Cheng
University of Electronic Science and Technology of China, Chengdu, China
Y
Yuchen Su
University of Electronic Science and Technology of China, Chengdu, China
Q
Qianxin Xia
University of Electronic Science and Technology of China, Chengdu, China
G
Guoming Lu
University of Electronic Science and Technology of China, Chengdu, China
J
Jielei Wang
University of Electronic Science and Technology of China, Chengdu, China
Wen Li
Wen Li
Data Intelligence Group, UESTC
Machine LearningComputer VisionDomain AdaptationTransfer LearningWeb Data