Learn from Whoever Is Right: Answer-Verified Multi-Teacher Distillation for Multi-Domain LLMs

📅 2026-09-02
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出MT-SDPO方法,通过验证答案确定可靠的教师模型,以解决多领域大语言模型能力整合难题。
📝 Abstract
Modern large language models (LLMs) rely on reinforcement learning to build strong capabilities in individual domains, but integrating those capabilities into a single deployable model remains challenging. By routing each sample to the teacher whose domain matches it, existing approaches let a domain label decide which teacher provides supervision. However, domain expertise holds only on average: the matched teacher is not always correct on a given sample, while a teacher from another domain sometimes is. The reliable teacher therefore has to be identified per sample, not per domain. In this paper, we introduce Multi-Teacher Self-Distillation Policy Optimization (MT-SDPO), an on-policy distillation method that unifies several frozen teachers into one student model. MT-SDPO consists of three components: (1) self-anchors, where a rollout is supervised by a correct rollout from its own group; (2) answer-verified eligibility, where a teacher may supervise a sample only if its own answer passes a verifier; and (3) privileged distillation, which merges the anchor and all verified feedback into one context that an exponential moving average self-teacher reads and the student does not, thereby keeping one policy at deployment. Across five students from three model families, MT-SDPO lifts the weakest domain of Qwen3-8B by 14.79 points and narrows its domain gap by 74.7%, a better balance than serving one matched teacher per domain. Verified reliability, not domain membership, should decide who teaches. Code is available at https://github.com/hexixiang/MT-SDPO.
Problem

Research questions and friction points this paper is trying to address.

Multi-Teacher Distillation
Large Language Models
Domain Expertise
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-Teacher Self-Distillation Policy Optimization
answer-verified eligibility
privileged distillation
🔎 Similar Papers
No similar papers found.
X
Xixiang He
National University of Defense Technology
X
Xingming Li
National University of Defense Technology
B
Baiqi Wu
Zhejiang University
Qiyao Sun
Qiyao Sun
QueenMary University of London
AI Scientist
X
Xuanyu Ji
National University of Defense Technology
A
Ao Cheng
National University of Defense Technology
Qingyong Hu
Qingyong Hu
Ph.D. of Computer Science, University of Oxford
3D VisionPhotogrammetryPoint Cloud ProcessingAutonomous Driving