Benchmarking LLM Competence on Logical Inference over Probability Operators
This study addresses the difficulty large language models (LLMs) face in distinguishing genuine logical reasoning from superficial pattern matching when processing probabilistic expressions such as “might” or “must,” revealing a fundamental deficiency in their capacity for logical inference under uncertainty. The authors present the first systematic benchmark comprising 14,320 synthetically generated English samples across 15 reasoning templates, carefully controlling variables including question formulation, negation strategies, and surface-level content to isolate models’ logical capabilities with modal expressions. Introducing a “reasoning floor” metric—defined as the worst-case accuracy on Yes/No tasks—the work evaluates whether models exhibit true reasoning rather than response biases. Evaluation of 29 prominent LLMs shows only nine surpass random performance, with most exhibiting significant, logic-irrelevant deviations tied to question phrasing, verb animacy, and named entity attributes.