Optimizing Sparse Outcomes Through Dense Behavioral Signals via Value-Guided Preference Distillation

📅 2026-09-13
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
通过多目标强化学习和密集行为信号优化长期对话结果,利用偏好蒸馏方法提高用户留存率和积极行为。
📝 Abstract
Aligning multi-turn dialogue agents is usually framed as matching turn-level human preferences, yet direct optimization of long-term outcomes is often ineffective and prone to reward hacking. We formulate long-horizon dialogue optimization as a multi-objective reinforcement learning problem and train a multi-head value model that predicts a vector of observed user behaviors across multiple look-ahead horizons. Our findings demonstrate that a scalarized composite of dense auxiliary behavioral signals enables effective credit assignment and optimization of sparse outcomes. However, optimizing unconstrained single-objective proxies might induce policy degradations that are harmful when the agent is exposed to real users. To identify these failure modes prior to deployment, we establish a safety framework combining counterfactual user simulation with a validated dialogue-level outcome model to evaluate preference weightings and policy optimization methods. Finally, we demonstrate that distilling multi-objective value preferences into the policy via reference-anchored preference optimization matches on-policy online RL at a small fraction of its compute budget. Live A/B testing confirms that our distilled policy significantly improves long-term user retention, while simultaneously enhancing the positive behaviors and therapeutic-process markers.
Problem

Research questions and friction points this paper is trying to address.

multi-turn dialogue
long-term outcomes
reward hacking
policy degradation
user retention
Innovation

Methods, ideas, or system contributions that make the work stand out.

value-guided preference distillation
multi-objective reinforcement learning
dense auxiliary behavioral signals
safety framework with counterfactual user simulation
reference-anchored preference optimization
🔎 Similar Papers
No similar papers found.
Ziyi Zhu
Ziyi Zhu
Intel, Columbia University
Optical Interconnects
D
Daniel R. Cahn
Slingshot AI
T
Thomas D. Hull
Slingshot AI
C
Caitlin A. Stamatis
Slingshot AI
O
Olivier Tieleman
Slingshot AI
G
Guilherme B. Freire
Slingshot AI
Jinghong Chen
Jinghong Chen
University of Cambridge
Natural Language ProcessingDialogue System