TrustDABench: Benchmarking Reliability and Robustness of LLMs for Structured Data Analysis

📅 2026-08-25
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过创建TrustDABench基准,利用19种扰动操作评估了八个大模型在结构化数据分析中的可靠性和鲁棒性问题。
📝 Abstract
LLMs are increasingly used to analyze spreadsheets, CSV files, and other structured data, but producing a correct-looking answer is not the same as producing a trustworthy analysis. A trustworthy result should be supported by a valid path from the user question to the relevant data evidence. This requirement creates two diagnostic questions: whether an LLM can refuse to answer or ask for clarification when such a path does not exist, and whether it can preserve the correct analysis when the same evidence is expressed in different table forms. We introduce TrustDABench, a benchmark that operationalizes these questions as reliability and robustness. Starting from the evidence-path view, we derive 19 perturbation operators and instantiate them through an Agentic-LLM-based generation framework. TrustDABench contains 2,340 human-verified perturbed instances, and we evaluate eight representative LLMs. The results show substantial headroom: the best reliability result is only 24.21% average MRS, achieved by GPT-5.5, while the best robustness result still has 9.10% average ASR, achieved by Claude-Sonnet-5. The failures are systematic: models rarely detect conflicting evidence, often continue along executable but unsupported analysis paths, and remain sensitive to perturbations that change observation boundaries or cross-table relations. These findings suggest that stronger evidence-boundary recognition and representation-invariant reasoning are still needed for reliable structured-data analysis.
Problem

Research questions and friction points this paper is trying to address.

reliability
robustness
structured data analysis
evidence path
perturbation
Innovation

Methods, ideas, or system contributions that make the work stand out.

TrustDABench
reliability and robustness
structured data analysis
evidence-path view
perturbation operators
🔎 Similar Papers
No similar papers found.
Boshen Shi
Boshen Shi
中移九天人工智能研究院
Graph Neural NetworksTransfer LearningTable Mining
Y
Yize Liu
School of Computer and Cyberspace Security, Communication University of China
C
Chen Zhao
China Mobile Jiutian Artificial Intelligence Technology (Beijing) Co., Ltd.
C
Ce Chi
China Mobile Jiutian Artificial Intelligence Technology (Beijing) Co., Ltd.
Zhendong Wang
Zhendong Wang
University of Science and Technology of China (USTC)
Computer VisionDeep LearningGenerative ModelAIGC
X
Xing Wang
China Mobile Jiutian Artificial Intelligence Technology (Beijing) Co., Ltd.
Junlan Feng
Junlan Feng
Chief Scientist at China Mobile Research
Natural LanguageMachine LearningSpeech ProcessingData Mining