DI-Bench: Systematically Generating In-Domain Data Intelligence Benchmarks for Enterprise Agents

📅 2026-09-04
📈 Citations: 0
Influential: 0
📄 PDF
📝 Abstract
Evaluating enterprise agents on domain-specific benchmarks is critical, yet public benchmarks rarely evaluate whether agents can integrate business knowledge with analytical computation, and constructing such benchmarks manually is costly. We present DI-Bench, a pipeline for generating realistic benchmarks for data intelligence (DI), the practice of extracting insights from large volumes of enterprise data. To emulate realistic DI tasks that require both computation and knowledge retrieval, DI-Bench builds an artifact linkage graph over data tables, dimensions, metrics, and documents to form questions involving structured data and associated knowledge. Ground truth answers are derived via query execution, followed by LLM question generation and validation. Applied to two public datasets, the pipeline produces a 731-task benchmark covering knowledge retrieval, analytical computation, and rule-grounded reasoning. To show the discriminatory capability and difficulty of the benchmark, we evaluate four models, revealing a substantial finding: models achieve only 32% accuracy when doing computational tasks where retrieved business rules modify the computation.
Problem

Research questions and friction points this paper is trying to address.

enterprise agents
domain-specific benchmarks
business knowledge
analytical computation
Innovation

Methods, ideas, or system contributions that make the work stand out.

DI-Bench
data intelligence
artifact linkage graph
LLM question generation
business knowledge integration
J
Jiangyun Zhang
Amazon.com
K
Kristen Surrao
Amazon.com
T
Torpong Nitayanont
Amazon.com
Y
Yupei Zhang
Amazon.com
R
Roopali Singh
Amazon.com
Zhiyu Chen
Zhiyu Chen
Amazon
Conversational AILarge Language ModelsInformation RetrievalNatural language Processing
J
Julia Huang
Amazon.com
Z
Zhou Tang
Amazon.com
S
Shayan Ali Akbar
Amazon.com
Omar Alonso
Omar Alonso
Amazon
Information RetrievalEvaluationLabelingKnowledge graphs
Erwin Cornejo
Erwin Cornejo
Amazon
AIDeep LearningNatural Language Processing
Y
Yuan Li
Amazon.com
Y
Yi Zhang
Amazon.com