Benchmarking Language Models for Statistical Problem Formulation

📅 2026-09-01
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文针对统计问题表述这一上游步骤,通过构建StatFormBench基准来评估大型语言模型在统计任务识别及变量确定上的表现。
📝 Abstract
Large language models (LLMs) are increasingly used as assistants for statistical and data science work, yet existing evaluations largely assume the analysis target is already specified. In practice, users arrive with informal goals and heterogeneous data, leaving the model to decide what statistical task is implied and which data are relevant. We first formalize this upstream step as Statistical Problem Formulation and decompose it into two subtasks: (1) Statistical Problem Classification and (2) Variable Identification & Role Assignment. We then introduce StatFormBench, a benchmark built from five cross-domain statistics textbooks and a data science case library, covering diverse problem types, data representations, and scenario styles. It contains 1,013 samples spanning 20 coarse-grained and 85 fine-grained statistical problem categories. Across 14 open- and closed-source LLMs, the best zero-shot models reach only 72.0 fine-grained classification accuracy and 63.2 variable set overlap. No model performs consistently best across the two subtasks, while enhanced prompting strategies yield only limited or inconsistent gains. We release the benchmark data on Hugging Face at https://huggingface.co/datasets/THU-CongLab/StatFormBench and the evaluation code on GitHub at https://github.com/THU-CongLab/StatFormBench.
Problem

Research questions and friction points this paper is trying to address.

Large language models
Statistical Problem Formulation
Variable Identification
Role Assignment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Statistical Problem Formulation
Variable Identification & Role Assignment
StatFormBench
🔎 Similar Papers
No similar papers found.
C
Chen Wang
Department of Statistics and Data Science, Tsinghua University
J
Junzhe Zhao
Department of Statistics and Data Science, Tsinghua University
Xin Cong
Xin Cong
Tsinghua University
Tool LearningAutonomous AgentLarge Language ModelKnowledge Graph
W
Wanlu Deng
Department of Statistics and Data Science, Tsinghua University
Ke Deng
Ke Deng
Tsinghua University
StatisticsData Science