MINT: Min-Selection Preference Distillation for Balanced Multi-Objective Alignment
This study addresses the optimization imbalance caused by reward aggregation in multi-objective preference training by proposing MINT. Grounded in the limit theory of generalized means, this method transforms traditional weighted summation into weakest-target ranking. By integrating minimal selection preference distillation with Direct Preference Optimization (DPO), MINT generalizes from weighted sums to worst-case selection with merely a single-line code modification. Experimental results demonstrate that MINT significantly improves scores for underperforming objectives while reducing performance disparity. Notably, it surpasses human expert performance in emotional support tasks, offering an efficient solution for balanced multi-objective alignment.