VIBE-Bench: Evaluating Personalized Large Language Models When Profiles Don't Mean Preferences

📅 2026-09-01
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究解决了个性化大语言模型中用户档案与偏好概念不一致的问题,通过引入VIBE-Bench基准来评估模型跨概念偏好的推理能力。
📝 Abstract
Personalized Large Language Models (PLLMs) aim to tailor responses to individual users, where a central challenge is preference reasoning: inferring query-relevant preferences from user-related history. Existing benchmarks, however, largely assume that such preference can be retrieved from semantically related history. We study an underexplored but practically important regime, profile-preference conceptual misalignment (PRCM), where observable profile cues and query-specific preferences lie in different concept spaces, making semantic retrieval inconsistent for personalization. We introduce VIBE-Bench, a benchmark with two psychology-grounded tasks, 3,504 personas and 12,239 dialogues, including a manually verified gold test set, and requires cross-concept preference reasoning beyond surface semantic overlap. Experiments with several personalization methods show that current PLLMs largely rely on shallow semantic correlations and fail to acquire robust cross-concept mappings. These findings establish PRCM as a distinct failure regime in PLLMs and position VIBE-Bench as a focused testbed for advancing preference reasoning beyond semantic matching.
Problem

Research questions and friction points this paper is trying to address.

Personalized Large Language Models
Preference Reasoning
Profile-Preference Conceptual Misalignment
Semantic Retrieval
Cross-Concept Preference Reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

VIBE-Bench
Profile-Preference Conceptual Misalignment (PRCM)
Cross-concept Preference Reasoning
🔎 Similar Papers