Verifiable Social Reasoning for LLM Assistants

📅 2026-09-15
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究提出Fuse框架,通过模拟多代理交互来评估LLM助手的社会推理能力,解决社会情境主观性及缺乏可验证事实的问题。
📝 Abstract
LLM assistants are widely used for daily social advice, yet evaluating their social reasoning in such consultation settings remains challenging since (i) it requires setups where the assistant learns about social situations from subjective user narratives, and (ii) social properties, such as others' intentions, typically lack verifiable ground truth. To address these challenges, we introduce Fuse, a multi-agent simulation framework for studying user-mediated social reasoning. In Fuse, a target agent with a hidden motive interacts with other agents including one representing the user, who then consults the evaluated assistant to infer the target's motive, providing verifiable ground truth by construction. Simulation faithfulness is validated through a human study with 24k annotations. We apply Fuse to 12 LLMs and demonstrate its analytical utility by systematically isolating key factors, showing that (i) user mediation compounds the inherent difficulty of social reasoning; (ii) LLMs exhibit systematic sensitivity to biased user framing; (iii) models can require more details than humans need to reach a correct prediction; and (iv) longer conversations do not always improve performance despite providing opportunities for clarifying questions. We open-source Fuse and a dataset with 21k examples.
Problem

Research questions and friction points this paper is trying to address.

LLM Assistants
Social Reasoning
Verifiable Ground Truth
User Narratives
Innovation

Methods, ideas, or system contributions that make the work stand out.

Verifiable Social Reasoning
Multi-agent Simulation
User-mediated
Large Language Models (LLMs)
Social Inference