Preserving General Capabilities during Domain Specialization with Uncertainty-Calibrated MOPD

📅 2026-08-27
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
该研究通过不确定性校准的MOPD方法解决领域特化过程中通用能力下降的问题,实验显示在保持垂直领域性能的同时提高了通用能力。
📝 Abstract
Specializing large language models to vertical domains improves domain-specific behavior but often degrades general capabilities such as reasoning, coding, instruction following, and creative writing. We study this domain--general trade-off in Multi-Teacher On-Policy Distillation (MOPD), where a specialized student is supervised on its own sampled trajectories by domain and general teachers. Standard MOPD faces two limitations: ordinary on-policy sampling rarely exposes tokens with large positive teacher--student advantages, while the advantage sign alone does not establish whether the resulting update direction is reliable. We propose uncertainty-calibrated MOPD to address these limitations. Dual-temperature sampling broadens the candidate trajectory pool, and positive-advantage-density filtering selects trajectories with stronger positive learning signals. Centered log-likelihood (CLL) filtering then computes an entropy-calibrated teacher-endorsement score and probabilistically retains token updates according to direction--endorsement consistency. Experiments on role-playing and medical-domain specialization show that our method improves the general-capability average over standard MOPD by $4.73\%$ and $10.84\%$, respectively, while maintaining vertical-domain performance. Ablations and diagnostic analyses further confirm that the gains do not merely result from a larger rollout budget and that the proposed trajectory- and token-level mechanisms address their intended failure modes.
Problem

Research questions and friction points this paper is trying to address.

domain specialization
general capabilities
language models
trade-off
Innovation

Methods, ideas, or system contributions that make the work stand out.

Uncertainty-Calibrated MOPD
Dual-Temperature Sampling
Positive-Advantage-Density Filtering
Centered Log-Likelihood (CLL) Filtering
🔎 Similar Papers
Ziyuan Liu
Ziyuan Liu
Unknown affiliation
RoboticsManipulation and GraspingComputer VisionMachine Learning
J
Jiao Ou
Kuaishou Technology, Beijing, China
Jian Liang
Jian Liang
Kuaishou Inc.
transfer learninggraph learning
R
Ruiming Tang
Kuaishou Technology, Beijing, China
C
Cheng Luo
Kuaishou Technology, Beijing, China