Do LLMs Trust the Accuser or the Accusation? Measuring Belief Shifts in Werewolf

📅 2026-09-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过在狼人杀游戏中测量AI模型信念变化,评估其交流技能,发现大型模型能更好地区分狼人和村民,但仍受指控影响。
📝 Abstract
Social-deduction games such as Werewolf are increasingly used to evaluate LLM agents, but existing evaluations often rely on final game outcomes. We propose a belief-shift evaluation benchmark in Werewolf for analyzing communication skills through belief updating. Using LLM-played games, we annotate suspicion and accusation messages and measure how an observing village-side model's beliefs change after each message. We evaluate 40 open-weight LLM configurations on 1,224 annotated messages. Our results show that larger models better distinguish true wolves from villagers based on game history, but accusations still strongly influence their beliefs. Models become more suspicious of the accused target and less suspicious of the accuser, especially when the accuser is trusted, even if the accuser is wolf-aligned. Larger models better resist accusations from accusers they already distrust. Overall, our findings suggest that current open-weight LLMs up to 120B parameters still struggle to integrate accusation content with source trust in strategic communication. Our benchmark and code are available at https://rlg.iis.sinica.edu.tw/papers/werewolf-accusation-benchmark.
Problem

Research questions and friction points this paper is trying to address.

Werewolf
belief shift
social-deduction games
LLM evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

belief-shift evaluation
social-deduction games
Werewolf
large language models
strategic communication
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yu-Yu Yang
Department of Computer Science, National Yang Ming Chiao Tung University, Taiwan
Ti-Rong Wu
Ti-Rong Wu
Institute of Information Science, Academia Sinica
Reinforcement learningPlanningComputer gamesDeep learningArtificial intelligence
H
Hung Guei
Institute of Information Science, Academia Sinica, Taiwan; College of Artificial Intelligence, National Yang Ming Chiao Tung University, Taiwan
H
Hsing-Yu Chen
Department of Computer Science, National Yang Ming Chiao Tung University, Taiwan; Institute of Information Science, Academia Sinica, Taiwan
I-Chen Wu
I-Chen Wu
National Chiao Tung University
computer gamesArtificial Intelligence