Benchmarking LLM Competence on Logical Inference over Probability Operators

📅 2026-07-29
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the difficulty large language models (LLMs) face in distinguishing genuine logical reasoning from superficial pattern matching when processing probabilistic expressions such as “might” or “must,” revealing a fundamental deficiency in their capacity for logical inference under uncertainty. The authors present the first systematic benchmark comprising 14,320 synthetically generated English samples across 15 reasoning templates, carefully controlling variables including question formulation, negation strategies, and surface-level content to isolate models’ logical capabilities with modal expressions. Introducing a “reasoning floor” metric—defined as the worst-case accuracy on Yes/No tasks—the work evaluates whether models exhibit true reasoning rather than response biases. Evaluation of 29 prominent LLMs shows only nine surpass random performance, with most exhibiting significant, logic-irrelevant deviations tied to question phrasing, verb animacy, and named entity attributes.
📝 Abstract
Both expressions of uncertainty and inferences are ubiquitous in natural language, and valid inferences over natural-language expressions of uncertainty are necessary for not only everyday conversations but also for high-stakes domains such as medicine and law. While large language models are increasingly evaluated on logical reasoning tasks, disentangling principled, symbolic reasoning from clever surface-level pattern matching is fraught with difficulty. We introduce a benchmark for reasoning over probability operators--inference over sentences with gradable epistemic modals (e.g., probably, might, must) containing 14,320 procedurally-generated English prompts across fifteen inference templates, systematically varying question form, negation strategy, and surface content. Evaluating 29 models, we find that most show answer biases independent of the logical form, a systematic preference for Yes or No. We summarize this with a competence floor: the worse of a model's accuracy on Yes-correct and No-correct items. Only 9 of 29 models exceed random chance. We also test variations in question form, verb phrases/activity, and both the gender and origin of names used in the prompts, finding biases across every axis.
Problem

Research questions and friction points this paper is trying to address.

logical inference
probability operators
epistemic modals
answer bias
large language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

probability operators
logical inference
epistemic modals
competence floor
procedurally-generated benchmark
💼 Related Jobs
No related jobs found.