Toward Robust Personalized Alignment for LLMs: Mitigating Persona Drift in Multi-Turn Dialogue

📅 2026-09-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
该研究针对个性化语言模型中的人格漂移问题,提出CORE方法,通过区分局部证据和持久人格状态并进行有选择性的更新,以提高多轮对话中的个性化对齐和鲁棒性。
📝 Abstract
Persona drift remains a central challenge for personalized language models, as user profiles evolve over long interactions rather than remain permanently fixed. Models must therefore revise persistent persona states when preferences genuinely change, while avoiding updates driven by transient, ambiguous, or unresolved observations. We propose CORE, which separates turn-local evidence from persistent persona-state revision and selectively updates grounded user preferences through uncertainty-aware belief revision. We also introduce PERSIST, a held-out post-anchor benchmark for persona-state robustness under sequential interaction stress, covering ambiguity, conflict, and controlled social influence. Across ALOE, PersonaChat, and PERSIST, CORE improves personalized alignment and robustness, with complementary gains in normalized closed-slot state fidelity. Human evaluation and mechanistic controls further support explicit update control beyond stronger generation or persistent memory alone.
Problem

Research questions and friction points this paper is trying to address.

Persona Drift
Personalized Language Models
Multi-Turn Dialogue
Persistent Persona States
User Preferences
Innovation

Methods, ideas, or system contributions that make the work stand out.

CORE
Uncertainty-aware Belief Revision
Persona Drift
PERSIST Benchmark
Sequential Interaction Stress
🔎 Similar Papers
No similar papers found.