MIRA-Math: A Benchmark for Minimal Information Requesting and Mathematical Reasoning

📅 2026-07-08
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses a critical gap in existing mathematical reasoning benchmarks, which typically provide complete information and thus fail to evaluate a model’s ability to proactively request missing facts. The authors introduce MIRA-Math, the first benchmark to formalize and assess the diagnostic capability of “minimal information requesting”: given a math problem missing exactly one atomic fact, models must precisely query for the absent information in natural language and, under a strict budget, integrate the retrieved response to produce an exact answer. Through deterministically generated instances, typed prompting protocols, and constrained LLM response channels—augmented by verification and answer-checking mechanisms—the benchmark ensures reproducible evaluation while decoupling information-seeking from reasoning. Experiments across 2,310 instances spanning nine mathematical domains reveal a dissociation between state-of-the-art and smaller models’ success in information requesting versus final answer accuracy, effectively pinpointing critical failure modes in reasoning chains.
📝 Abstract
Mathematical reasoning benchmarks typically provide all facts needed to solve each problem, while interactive benchmarks often mix reasoning with tools, retrieval, and long-horizon dialogue. We introduce MIRA-Math, a benchmark for a narrower diagnostic capability: solving mathematical problems whose full latent state has a unique answer, but whose solver-facing view is missing exactly one necessary atomic fact. The solver must request the missing information in natural language under a strict budget and then integrate the returned fact into an exact final answer. A fixed constrained LLM responder sees only the dataset-provided atomic fact and must either offer the quoted fact when the request matches it, or decline otherwise. Thus, instance generation, typed hint specifications, validation, and final-answer verification are deterministic, while request metrics are measured under a fixed LLM-mediated responder channel. MIRA-Math contains 2{,}310 generated instances from 22 typed mathematical families spanning algebra, probability, linear systems, discrete structures, signal processing, Markov chains, circuits, interpolation, and numerical boundary-value problems. Experiments across frontier and small models show that request success and final-answer accuracy are separable: models may ask for the right fact yet fail the downstream computation, or fail before obtaining the canonical hint. We release generators, verifiers, prompts, run metadata, and dataset documentation to support reproducible evaluation of minimal information requesting in mathematical reasoning.
Problem

Research questions and friction points this paper is trying to address.

mathematical reasoning
information requesting
benchmark
missing fact
atomic fact
Innovation

Methods, ideas, or system contributions that make the work stand out.

minimal information requesting
mathematical reasoning
constrained LLM responder
atomic fact retrieval
deterministic benchmark
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
C
Charbel Al Bateh
Department of Electrical and Computer Engineering, Lebanese American University, Byblos, Lebanon
S
Samer Saab Jr.
Department of Electrical and Computer Engineering, Lebanese American University, Byblos, Lebanon