Evaluation of Deontic Conditional Reasoning in Large Language Models: The Case of Wason's Selection Task

📅 2026-03-06
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study investigates the performance of large language models (LLMs) in deontic conditional reasoning and examines their similarities to and differences from human cognitive biases. Addressing the lack of systematic distinction between deontic and descriptive rules in prior work, we introduce deontic modality into the Wason selection task for the first time, constructing a structured dataset to enable controlled experimentation. Drawing on theories of confirmation bias and matching bias, we conduct systematic behavioral experiments with mainstream LLMs. Results show that models perform significantly better under deontic rules than under descriptive ones, and their error patterns exhibit a matching bias akin to that observed in human reasoning. These findings reveal that LLMs’ reasoning is domain-sensitive and susceptible to cognitive biases analogous to those affecting human judgment.

Technology Category

Application Category

📝 Abstract
As large language models (LLMs) advance in linguistic competence, their reasoning abilities are gaining increasing attention. In humans, reasoning often performs well in domain specific settings, particularly in normative rather than purely formal contexts. Although prior studies have compared LLM and human reasoning, the domain specificity of LLM reasoning remains underexplored. In this study, we introduce a new Wason Selection Task dataset that explicitly encodes deontic modality to systematically distinguish deontic from descriptive conditionals, and use it to examine LLMs'conditional reasoning under deontic rules. We further analyze whether observed error patterns are better explained by confirmation bias (a tendency to seek rule-supporting evidence) or by matching bias (a tendency to ignore negation and select items that lexically match elements of the rule). Results show that, like humans, LLMs reason better with deontic rules and display matching-bias-like errors. Together, these findings suggest that the performance of LLMs varies systematically across rule types and that their error patterns can parallel well-known human biases in this paradigm.
Problem

Research questions and friction points this paper is trying to address.

deontic reasoning
Wason Selection Task
large language models
conditional reasoning
cognitive bias
Innovation

Methods, ideas, or system contributions that make the work stand out.

deontic reasoning
Wason selection task
large language models
matching bias
conditional reasoning
H
Hirohiko Abe
Keio University, Tokyo, Japan
K
Kentaro Ozeki
Keio University, Tokyo, Japan; University of Tokyo, Tokyo, Japan
R
Risako Ando
Keio University, Tokyo, Japan
T
Takanobu Morishita
Keio University, Tokyo, Japan
Koji Mineshima
Koji Mineshima
Keio University
M
Mitsuhiro Okada
Keio University, Tokyo, Japan