MINT: Min-Selection Preference Distillation for Balanced Multi-Objective Alignment

📅 2026-08-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the optimization imbalance caused by reward aggregation in multi-objective preference training by proposing MINT. Grounded in the limit theory of generalized means, this method transforms traditional weighted summation into weakest-target ranking. By integrating minimal selection preference distillation with Direct Preference Optimization (DPO), MINT generalizes from weighted sums to worst-case selection with merely a single-line code modification. Experimental results demonstrate that MINT significantly improves scores for underperforming objectives while reducing performance disparity. Notably, it surpasses human expert performance in emotional support tasks, offering an efficient solution for balanced multi-objective alignment.
📝 Abstract
Aligning a language agent to several objectives at once is a persistent failure mode of preference-based training: when objectives are combined additively, optimization collapses onto whichever is cheapest to improve and sacrifices the rest, so a support agent learns to sound warm while giving no real help. The root issue is that an additive reward has no notion of balance. We introduce Mint (MIN-selection preference disTillation), a one-line change to preference distillation: rather than ranking sampled candidates by a weighted sum of rewards, we rank them by their weakest objective, distilling the best-balanced candidate over the most lopsided one with an unchanged DPO objective. This is the p -> negative infinity limit of a generalized-mean family spanning additive to worst-case selection. Across cooperative emotional support and adversarial negotiation, min-selection lifts both objectives while sharply cutting their imbalance; on emotional support it raises the weaker axis from 0.37 to 0.64 (p < 10^-40), surpassing human experts and persisting across full multi-turn rollouts. A turn-by-turn analysis yields our central finding: min-selection corrects imbalance in proportion to how imbalanced the reference policy is, and its benefit endures over an interaction precisely as long as that imbalance does.
Problem

Research questions and friction points this paper is trying to address.

Multi-Objective Alignment
Preference-based Training
Optimization Collapse
Reward Imbalance
Language Agent
Innovation

Methods, ideas, or system contributions that make the work stand out.

Min-Selection Preference Distillation
Multi-Objective Alignment
Balanced Optimization
DPO
Generalized Mean
🔎 Similar Papers
No similar papers found.
T
Tony Tu
Georgia Institute of Technology, Zillow Group
S
Sayan Chakraborty
Zillow Group
R
Ruomeng Xu
Zillow Group
T
Tony Qin
Zillow Group
A
Austin Tian
Zillow Group