Level-k Distinguishable Mechanisms for Evaluating Bounded Rationality in LLMs

📅 2026-08-21
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为评估大型语言模型在有限理性环境中的策略推理深度,研究者构建了新的游戏结构,并通过递归推理和对手游戏数据归纳来评估模型的战略深度。
📝 Abstract
Strategic depth of reasoning is essential for human interaction of Large Language Models (LLMs) operating in boundedly rational environments. However, existing evaluations are primarily based on canonical games prevalent in pretraining corpora, making it difficult to disentangle true strategic reasoning from memorisation. To address this, we formalise a necessary level-K distinguishability condition for strategic depth inference and construct a suite of novel game structures that meet this standard. Using these games, we evaluate strategic depth in LLMs from both the Chain-of-Thought tokens and actual actions under recursive reasoning and an inductive trace of opponent game-play data. Across experimental trials spanning four LLMs, four game structures, and ten levels of iterated reasoning, we find that model models maintain accurate strategic depth under recursive reasoning, with strong internal consistency between stated reasoning and actions at every level. Errors arise from using the wrong number of iterated depth of reasoning steps, not from computing best responses incorrectly. However, inductive inference from opponent play degrades accuracy sharply and unevenly across games, and explicit strategic mentalizing in the chain of thought substantially improves overall performance.
Problem

Research questions and friction points this paper is trying to address.

strategic reasoning
bounded rationality
large language models
level-k distinguishability
Innovation

Methods, ideas, or system contributions that make the work stand out.

level-K distinguishability
strategic depth inference
recursive reasoning
inductive inference
strategic mentalizing
🔎 Similar Papers