Learning from Consensus and Disagreement: Unsupervised On-Policy Self-Distillation with Minority-Trajectory Contrast

πŸ“… 2026-08-09
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses the limitation of existing policy-based self-distillation methods, which rely on external supervision and struggle to enhance language models’ reasoning capabilities under fully unsupervised conditions. To overcome this, we propose CoDA, a novel framework that, for the first time, constructs reliable privileged information in an unsupervised setting by leveraging consensus among model-generated reasoning trajectories as positive guidance and minority trajectories as negative samples for contrastive calibration. CoDA further integrates a frozen self-teacher, answer-level consensus identification, and a KTO-style reference anchoring mechanism to effectively mitigate error amplification caused by incorrect consensus. Experimental results demonstrate that CoDA significantly outperforms self-generation baselines on competitive mathematical reasoning benchmarks, simultaneously improving both reasoning performance and training stability.
πŸ“ Abstract
On-policy self-distillation improves language-model reasoning by querying a teacher on states actually visited by the student. Recent methods create a powerful information asymmetry by exposing the teacher to privileged context, yet they fundamentally rely on external supervision---such as gold solutions or verifiers---to construct this advantage. We introduce CoDA (Consensus and Disagreement Alignment), a fully unsupervised framework that creates reliable privileged information entirely from the latent uncertainty structure of a model's own unlabeled rollouts. CoDA extracts two complementary signals. In the positive branch, answer-level consensus identifies a stable reasoning mode, which conditions a frozen self-teacher to provide dense distributional guidance on fresh student trajectories. However, because agreement does not guarantee correctness, positive-only distillation risks amplifying correlated errors into a false consensus. To break this harmful feedback loop, CoDA incorporates a negative branch that exploits disagreement: minority trajectories are treated as unstable alternatives and gently penalized via a reference-anchored, KTO-style calibration objective. This unpaired binary feedback provides robust regularization without requiring the strong assumption that the consensus is the absolute ground truth. Empirical evaluations on competition-level mathematical benchmarks demonstrate that CoDA significantly improves reasoning, outperforming self-generated baselines and effectively stabilizing training against erroneous consensus.
Problem

Research questions and friction points this paper is trying to address.

unsupervised
self-distillation
reasoning
consensus
disagreement
Innovation

Methods, ideas, or system contributions that make the work stand out.

unsupervised self-distillation
on-policy learning
consensus-disagreement alignment
minority-trajectory contrast
reasoning stabilization
πŸ”Ž Similar Papers
2024-07-21arXiv.orgCitations: 1
J
Jiaxin Guo
School of Intelligence Science and Technology, Peking University; State Key Laboratory of General Artificial Intelligence, Peking University, Beijing, China
Y
Yanwei Yue
School of Intelligence Science and Technology, Peking University; State Key Laboratory of General Artificial Intelligence, Peking University, Beijing, China
X
Xuanbo Fan
School of Intelligence Science and Technology, Peking University; State Key Laboratory of General Artificial Intelligence, Peking University, Beijing, China
C
Chunyu Yang
Ucap Cloud
Yan Zhang
Yan Zhang
Peking University
Data MiningInformation RetrievalSocial NetworkNLP