Direct Diversity Optimization for Diverse Successful Trajectories in Preference Post-Training

📅 2026-09-09
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决LLM代理在序列决策任务中多样成功策略覆盖问题,提出直接多样性优化方法(DDO),结合分歧树收集与参考相对目标概率目标进行离线后训练。
📝 Abstract
LLM agents for sequential decision tasks are often post-trained with trajectory-level outcome labels, but such labels provide little supervision for preserving multiple successful branches from the same decision state. We study this problem as successful strategy coverage: how broadly a model realizes distinct successful strategies under a fixed rollout budget. We present Direct Diversity Optimization (DDO), an offline post-training method that combines Divergence-Tree Collection (DTC) with the Reference-Relative Target-Odds Objective (RTO). DTC constructs state-aligned branch sets rooted at shared decision states, and RTO trains the model to match reference-relative targets over successful alternatives. DDO achieves the strongest task success and successful strategy coverage among the compared post-training methods across BabyAI, BabaIsAI, and WebShop. It also achieves the highest recovery rate after local action replacement and higher task success and coverage than successful-only imitation and decoding-time diversification controls.
Problem

Research questions and friction points this paper is trying to address.

trajectory-level outcome labels
successful strategy coverage
sequential decision tasks
post-training
diversity optimization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Direct Diversity Optimization (DDO)
Divergence-Tree Collection (DTC)
Reference-Relative Target-Odds Objective (RTO)
🔎 Similar Papers
No similar papers found.
J
Junwon Ko
School of Electrical Engineering, KAIST
D
Dong-Jae Lee
School of Electrical Engineering, KAIST
M
Minchan Kwon
School of Electrical Engineering, KAIST
S
Sunghyun Baek
School of Electrical Engineering, KAIST
Junmo Kim
Junmo Kim
School of Electrical Engineering, KAIST
Statistical Signal ProcessingImage ProcessingComputer VisionMachine LearningInformation Theory