Investigating Assistant Bias in LLM User Simulators Using a Role Vector

📅 2026-08-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究使用角色向量方法分析并解决LLM用户模拟器中的助手偏置问题,通过模型激活对比用户与助手视角来提取用户角色向量。
📝 Abstract
LLM-based user simulators are increasingly used to evaluate autonomous agents at scale, in place of costly human evaluations. Despite this promise, these simulators exhibit "assistant bias," a tendency to cooperate and pursue task goals. They rarely reproduce the frustration or disengagement that real users exhibit, compromising evaluation validity. Prior work outlines that this bias is baked in during model training, which role-playing prompts fail to override. We analyze this bias from model activations, extracting a user role vector by contrasting how the model represents user versus assistant perspectives on the same dialogue. We observe two findings: (i) the user direction is identifiable in activations, elicits user-like behaviors, and captures characteristics distinct from assistant traits; and (ii) although user-role activation associates with simulation realism and steering strengthens it, it can exaggerate user behaviors and override individual user profiles. Together, our findings provide a representation-level analysis of LLM user simulators, confirming that assistant bias is structurally identifiable and that user behavior can be directionally analyzed.
Problem

Research questions and friction points this paper is trying to address.

Assistant Bias
User Simulators
Evaluation Validity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Assistant Bias
Role Vector
Model Activations
💼 Related Jobs
No related jobs found.
D
Daeheon Jeong
KAIST
Yoonjoo Lee
Yoonjoo Lee
KAIST
Human Computer InteractionNatural Language Processing
E
Eugene Choi
Seoul National University
S
Sinie van der Ben
ETH Zürich
J
Juho Kim
KAIST, SkillBench