AcCoRD: Evaluating User-Agent Collaboration Under Realistic User Preference Dynamics

📅 2026-08-27
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出AcCoRD基准,用于评估用户-代理协作中处理动态用户偏好的能力,通过在线购物和旅行规划两个领域测试五种前沿LLM,揭示了模型在应对偏好变化上的不足。
📝 Abstract
User preferences in user-agent collaboration are rarely static and fully-specified upfront: preferences are formed, revealed, adjusted, and relaxed during interaction. Existing benchmarks for evaluating user-agent collaboration focus almost exclusively on resolving underspecified preferences, thereby failing to capture the richer dynamics of real-world interaction. We introduce AcCoRD, a user-agent collaboration benchmark requiring agents to handle diverse user preference dynamics in two domains: online shopping and travel planning. We evaluate five frontier LLMs under two prompting strategies: vanilla ReAct, and an uncertainty-guided variant that prompts models to identify and resolve ambiguity about user preferences. Our results reveal that frontier models can handle underspecification but struggle to satisfy preferences that emerge or evolve mid-interaction and require more sophisticated uncertainty modeling. Further, prompting alone fails to elicit the required uncertainty recognition. We release AcCoRD as a resource for developing agents that can navigate the full complexity of real-world user preferences.
Problem

Research questions and friction points this paper is trying to address.

user-agent collaboration
preference dynamics
real-world interaction
underspecified preferences
Innovation

Methods, ideas, or system contributions that make the work stand out.

User-Agent Collaboration
Preference Dynamics
Uncertainty Modeling
Real-World Interaction
🔎 Similar Papers
No similar papers found.