From Atomic to Agentic: Towards Interpretable Evaluation of LLMs' Agentic Mathematical Capabilities

📅 2026-08-27
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为了解决现有数学基准仅评估最终答案的问题,本文提出了一种过程级基准,通过结构化分类和任务设计来评估大型语言模型的代理数学推理能力。
📝 Abstract
Large Language Models (LLMs) are evolving from performing end-to-end mathematical reasoning to integrating agentic intelligence. However, most existing math benchmarks evaluate only final answers. This outcome-oriented evaluation provides limited diagnostic value for identifying process-level failures or rigorous logic, failing to guide the transformation of LLMs into robust agents. To bridge this gap, we present a process-level benchmark designed to evaluate the inherent agentic mathematical reasoning abilities of LLMs. Our framework aligns problem-solving agentic behaviors with a structured taxonomy of reusable mathematical atomic capabilities. We design a comprehensive suite of planning, action, and feedback tasks across both textual and multimodal contexts, supported by an automated pipeline that synthesizes high-quality trajectories and produces fine-grained annotations via controlled LLM rewriting. Experiments reveal that models with similar end-to-end accuracy can exhibit markedly different agentic capability profiles. This demonstrates that process-level evaluation is crucial for interpreting the true potential of LLMs and guiding the development of next-generation mathematical agents.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
agentic intelligence
mathematical reasoning
process-level evaluation
benchmark
Innovation

Methods, ideas, or system contributions that make the work stand out.

process-level benchmark
agentic mathematical reasoning
automated pipeline
fine-grained annotations
🔎 Similar Papers
No similar papers found.