EnterpriseRAG: Benchmarking LLM Instruction Adherence and Robustness under Non-Ideal Enterprise Retrieval

📅 2026-08-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Enterprise-scale Retrieval-Augmented Generation (RAG) systems often exhibit insufficient instruction-following capabilities in real-world scenarios due to retrieval noise, missing knowledge, and factual inconsistencies—failure modes inadequately captured by existing benchmarks. To address this gap, this work introduces a novel benchmark comprising 983 expert-validated samples across six domains, featuring the first evaluation paradigm that jointly incorporates multidimensional constraints and non-ideal retrieval conditions to systematically simulate three representative failure types. The study evaluates 13 state-of-the-art large language models and finds that only 26.8% of responses satisfy all specified constraints, revealing a systemic breakdown in holistic instruction adherence masked by high individual constraint satisfaction rates. These findings underscore the necessity of explicit context-aware protocols and calibrated judgment mechanisms, establishing a reproducible, fine-grained foundation for evaluating industrial-grade RAG systems.
📝 Abstract
Enterprise RAG deployments face a critical reliability gap: while LLMs satisfy 80% of individual constraints, only 26.8% of responses meet all requirements simultaneously, revealing a 57-point orchestration gap. Existing benchmarks assume clean retrieval with simple queries, failing to capture production conditions where noisy documents and multi-dimensional constraints coexist. We introduce EnterpriseRAG, a benchmark of 983 expert-validated samples across six domains that systematically simulates three failure modes absent from prior work: retrieval noise, knowledge gaps, and factual conflicts, coupled with complex instructions. Evaluation of 13 state-of-the-art LLMs reveals a severe instruction adherence collapse, where high per-constraint satisfaction masks low holistic compliance. Critical findings expose deep barriers under knowledge gaps and factual conflicts, even with reasoning-enhanced inference, indicating production RAG requires explicit context-aware protocols and calibrated judgment. EnterpriseRAG provides a reproducible foundation for measuring and closing these gaps, directly informing deployment decisions for enterprise-scale RAG systems. We will release the benchmark and evaluation framework upon publication.
Problem

Research questions and friction points this paper is trying to address.

Enterprise RAG
instruction adherence
retrieval noise
knowledge gaps
factual conflicts
Innovation

Methods, ideas, or system contributions that make the work stand out.

EnterpriseRAG
instruction adherence
retrieval noise
knowledge gaps
factual conflicts
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
H
Huiqi Miao
Jiutian Research, China Mobile, Beijing, China
X
Xinbao Sun
Jiutian Research, China Mobile, Beijing, China
Bo Wang
Bo Wang
Professor of Department of Engineering Mechanics, Dalian University of Technology, China
structural and multidisciplinary optimizationaerospace advanced materials and lightweight structurelarge structural experim
F
Fanyu Meng
Jiutian Research, China Mobile, Beijing, China
L
Lijun Mei
Jiutian Research, China Mobile, Beijing, China
N
Na Wu
Jiutian Research, China Mobile, Beijing, China
D
Di Jin
Jiutian Research, China Mobile, Beijing, China
C
Chao Deng
Jiutian Research, China Mobile, Beijing, China
Junlan Feng
Junlan Feng
Chief Scientist at China Mobile Research
Natural LanguageMachine LearningSpeech ProcessingData Mining