PLC-DPO: Posterior Label Correction in Noisy and Ambiguous Preference Optimization

📅 2026-08-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决直接偏好优化中因数据噪声和标签模糊导致的问题,提出PLC-DPO方法,通过校准策略参考边界来修正标签,提高学习效率与准确性。
📝 Abstract
Direct Preference Optimization (DPO) simplifies alignment through pairwise comparisons but assumes all observed preferences are reliable. Real data often violates this assumption, leading to reversed, weak, or ambiguous labels that cause harmful policy updates. To address this, we propose Posterior Label Correction DPO (PLC-DPO) to robustly optimize preferences by routing each pair's training signal as a clean, flip, or tie case. The key idea is to use the calibrated policy-reference margin as online evidence to take appropriate correction actions. This reframes noisy preference learning as actively correcting supervision direction and strength rather than merely filtering suspicious examples. Across 57 dataset-model-benchmark cells, PLC-DPO obtains the best mean win rate against DPO (60.5 vs. 55.5 for the next-best method). Injected-noise and tie stress tests, human disagreement analysis, and self-confirmation diagnostics further show that the routing remains stable and distinguishes flipped from weakly directional pairs.
Problem

Research questions and friction points this paper is trying to address.

Noisy Preferences
Ambiguous Preferences
Policy Optimization
Direct Preference Optimization
Label Correction
Innovation

Methods, ideas, or system contributions that make the work stand out.

Posterior Label Correction
Noisy Preference Learning
Calibrated Policy-Reference Margin
Active Supervision Correction
Robust Optimization
🔎 Similar Papers