Score
Transferring knowledge from one or more teacher models into a smaller or otherwise constrained student by matching outputs, features, or behaviours (including policies), while preserving privileged information or subtle signals and enabling efficient inference.
Conventional knowledge distillation tightly couples teacher and student architectures, resulting in poor cross-architecture generalization and prohibitive retraining costs for each new student. Method: We propose the Generalized Teacher Network (GTN), the first architecture-agnostic teacher framework that models the student pool as a weight-sharing supernet and employs a capacity-aware conditional mechanism to dynamically adapt the teacher to diverse student architectures. GTN jointly trains the teacher and students in a single, distillation-aware optimization pass. Contribution/Results: GTN eliminates the need for per-student teacher training; its overhead is amortized across the student pool. Evaluated on multi-architecture student pools, GTN consistently improves accuracy by 1.2–2.8% over baseline distillation methods, significantly enhancing deployment flexibility and computational efficiency.
Existing machine learning teaching frameworks overlook representational alignment between teachers and students, prioritizing only model accuracy. Method: We propose GRADE, a representation-alignment-driven pedagogical optimization framework. Leveraging controlled machine–machine and machine–human teaching experiments, we formally define and quantify the relationship between representational alignment and teaching utility, introducing the alignment-driven teaching utility curve. We further design GRADE-Match, a cross-modal teacher–student matching algorithm that optimizes representational adaptation. Contribution/Results: Experiments demonstrate that improved representational alignment significantly enhances student task accuracy—moderated by class size and representation diversity. In simulated teaching settings, GRADE-Match achieves an average 12.3% improvement in learning outcomes. GRADE establishes a novel paradigm for interpretable, optimization-aware intelligent teaching systems grounded in representational alignment principles.
In privileged imitation learning, students often fail to replicate teacher behaviors due to limited observational capabilities—stemming from a fundamental asymmetry: the teacher’s policy is not designed for the student’s partially observable setting. To address this, we propose a joint teacher-student training framework. First, we incorporate an action-divergence approximation term into the teacher’s reward function, theoretically grounded in performance bounds to mitigate imitation failure. Second, we introduce a supervised behavioral alignment step that explicitly constrains the teacher’s policy to be imitable by the student. Third, we optimize the entire system via vision-driven, end-to-end reinforcement learning. Evaluated on maze navigation, vision-guided quadrotor flight, and dexterous manipulation tasks, our approach yields substantial improvements in student policy performance, empirically validating the efficacy of enhancing teacher imitability.
This work addresses the inefficient utilization of “dark knowledge” in knowledge distillation due to teacher–student model capacity mismatch. We identify two empirical regularities in large-capacity teacher outputs: (i) low discriminability among non-ground-truth class probabilities, yet (ii) stable inter-class relative affinity relationships. Building on this, we establish the first quantitative link between teacher capacity and dark knowledge structure, proposing a novel paradigm that enhances the discriminability of non-ground-truth logits to mitigate capacity mismatch—moving beyond conventional reliance solely on teacher accuracy. Methodologically, we integrate logit softening with temperature calibration, an inter-class discrepancy enhancement module, and a multi-teacher contrastive distillation framework. Experiments on CIFAR-100 and ImageNet demonstrate significant performance gains for lightweight student networks, consistently outperforming state-of-the-art methods including FitNet and RKD. The approach proves robust across diverse teacher–student capacity configurations.
Existing meta-learning methods suffer from limited generalizability, often being confined to specific algorithms or requiring differentiability assumptions. This paper proposes a general reinforcement learning–driven meta-learning framework that trains a teacher policy to dynamically guide arbitrary student algorithms—without imposing structural or differentiability constraints on the student. Key contributions include: (i) the first unified pedagogical paradigm for meta-learning; (ii) a parameter-behavior encoder that implicitly infers the student’s internal parameter state from its input-output behavior; and (iii) a reward function grounded in learning progress. Experiments across supervised and reinforcement learning tasks demonstrate that our framework significantly outperforms baselines relying on heuristic rewards and handcrafted state representations, validating its broad generalizability and empirical effectiveness.
This work challenges the common practice in knowledge distillation of naively matching a teacher model’s absolute feature representations, which overlooks the fact that such representations are only equivalent up to orthogonal transformations and isotropic scaling. From a geometric perspective, the paper proposes a new paradigm centered on representation equivalence classes: the student should instead learn class-invariant structures of the teacher’s representations—such as Gram matrices, centered kernel alignment (CKA), or principal subspaces—or leverage coordinate alignment for effective supervision. This framework unifies feature matching, relational distillation, and grafting approaches, revealing that logit-level matching is ultimately key to capability transfer. Experiments on Qwen2.5 and Llama-3.1 demonstrate that high CKA similarity alone is insufficient for performance recovery, while successful grafting hinges on boundary overlap in the training data coverage, thereby validating the proposed theory.
This work addresses the distribution mismatch commonly faced by large language model agents in supervised fine-tuning, where training relies on complete teacher demonstrations while testing depends on student-generated contexts. The authors formulate online policy data construction as a budget allocation problem and propose replacing lengthy or costly filtered teacher trajectories with a small number of unfiltered, short-step teacher continuations, strategically injected into student-induced critical contexts. By systematically exploring the design space of rollout policies, switching time distributions, continuation lengths, and filtering rules—and incorporating a dual-cost model accounting for both teacher inference and supervision signal retention—the method demonstrates strong empirical performance on HotpotQA, ALFWorld, and Terminal-Bench-Dev. Notably, it matches or exceeds existing critical-context filtering baselines on the first two benchmarks at lower computational cost, indicating that a few well-placed teacher steps can substantially enhance training efficiency.
This work addresses the issue of systematic bias propagation in student-teacher learning, where directly matching the teacher’s outputs can inadvertently transfer its biases to the student. To mitigate this, the authors propose “Residuals as Teachers” (RaT), a novel approach that introduces residual learning into the student-teacher framework: the teacher estimates the residual between the student’s current prediction and the target, effectively guiding the student to mimic a proximal gradient optimization process. Theoretical analysis demonstrates that RaT achieves minimax optimal convergence rates under non-asymptotic excess risk, whereas conventional soft-target matching suffers from a persistent approximation error. Empirical evaluations on both synthetic data and the ImageNette covariate shift classification benchmark confirm that RaT substantially outperforms baseline methods and effectively suppresses bias propagation.
This work addresses the challenge of transferring reasoning capabilities from large language models to smaller ones without retraining or reliance on labeled data. It introduces the “Master Key Hypothesis,” positing that model abilities are encoded along specific directions within a low-dimensional latent subspace. Building on this insight, the authors propose the UNLOCK framework, which extracts these capability directions via activation contrast, aligns subspaces across models of different scales using low-rank linear transformations, and injects the identified directions during inference to unlock latent reasoning abilities in the target small model. Experiments demonstrate substantial performance gains: for instance, Qwen1.5-7B achieves a 12.1% accuracy improvement on MATH and AGIEval Math benchmarks, while Qwen3-14B-Base even surpasses its post-trained counterpart.
This work addresses the challenge of transferring knowledge from pretrained black-box models when input features reside in a space inconsistent with the model’s expected domain. The authors propose a two-stage neural network approach that decomposes the target regression function into transferable and non-transferable components. The transferable part is estimated via unsupervised alignment using unlabeled cross-space feature pairs, while the non-transferable component is learned from a small set of labeled data. This method achieves, for the first time, knowledge transfer from one or multiple black-box models to heterogeneous feature spaces. Theoretically, it yields a risk upper bound strictly tighter than that of minimax estimators relying solely on labeled data, and adapts automatically to diverse scenarios. Empirical results demonstrate significant improvements in prediction accuracy on both synthetic and real-world datasets.