Do Large Language Models Possess a Theory of Mind? A Comparative Evaluation Using the Strange Stories Paradigm

📅 2026-02-20
🏛️ arXiv.org
📈 Citations: 1
Influential: 0
📄 PDF
🤖 AI Summary
研究通过改编的故事测试法比较了五种大型语言模型与人类在心智理论能力上的表现,以探讨这些模型是否能从文本中推断出他人的信念、意图和情感。
📝 Abstract
The study explores whether current Large Language Models (LLMs) exhibit Theory of Mind (ToM) capabilities -- specifically, the ability to infer others'beliefs, intentions, and emotions from text. Given that LLMs are trained on language data without social embodiment or access to other manifestations of mental representations, their apparent social-cognitive reasoning raises key questions about the nature of their understanding. Are they capable of robust mental-state attribution indistinguishable from human ability in its output, or do their outputs merely reflect superficial pattern completion? To address this question, we tested five LLMs and compared their performance to that of human controls using an adapted version of a text-based tool widely used in human ToM research. The test involves answering questions about the beliefs, intentions, and emotions of story characters. The results revealed a performance gap between the models. Earlier and smaller models were strongly affected by the number of relevant inferential cues available and, to some extent, were also vulnerable to the presence of irrelevant or distracting information in the texts. In contrast, GPT-4o demonstrated high accuracy and strong robustness, performing comparably to humans even in the most challenging conditions. This work contributes to ongoing debates about the cognitive status of LLMs and the boundary between genuine understanding and statistical approximation.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Theory of Mind
mental-state attribution
social-cognitive reasoning
understanding
Innovation

Methods, ideas, or system contributions that make the work stand out.

Theory of Mind
Large Language Models
Strange Stories Paradigm
Robustness
Statistical Approximation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
A
Anna Babarczy
Institute of General and Hungarian Linguistics, Research Centre for Linguistics
András Lukács
András Lukács
Institute of Mathematics, Eötvös Loránd University
Data ScienceArtificial IntelligenceMachine LearningCombinatoricsGraph Theory
P
Péter Vedres
Department of Cognitive Science, Faculty of Natural Sciences, Budapest University of Technology and Economics
Z
Zétény Bujka
Department of Cognitive Science, Faculty of Natural Sciences, Budapest University of Technology and Economics