LexAgentHallu: A Hierarchical Benchmark for Profiling Hallucinations in Legal Agents

📅 2026-09-09
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决法律代理中因工具调用和推理错误导致的幻觉问题,通过构建含3414个实例的LexAgentHallu基准,采用专家参与四阶段流程及细化度量方法进行评估。
📝 Abstract
As large language models are increasingly deployed as tool-augmented legal agents, they introduce agentic hallucinations where tool-call and reasoning errors cascade into fabricated holdings and miscited authority. However, existing legal benchmarks evaluate only single-turn QA with outcome-level metrics, while agentic hallucination benchmarks lack legal-specific diagnostic capability. Neither answers to what extent and how a legal agent hallucinates along its trajectory. To address these limitations, we introduce LexAgentHallu, a legal agentic hallucination benchmark designed to evaluate to what extent and how legal agents fail along multi-step trajectories. Built through a four-stage expert-in-the-loop pipeline, LexAgentHallu contains 3414 instances across 17 legal categories and 6 task types. Each instance is annotated under a dual-layer hallucination taxonomy of 7 high-level categories and 27 fine-grained subclasses, covering both substantive errors and agent-procedural failures. We further design fine-grained metrics that quantify to what extent and localize how each failure occurs along an agent's execution path. Our evaluation across 18 proprietary and open-source agents uncovers a Right-Answer-Wrong-Reason effect and reveals that hallucination subclasses cluster rather than scatter, forming distinct agentic framework, legal task, and category profiles. These findings, invisible to outcome-level evaluation, validate the diagnostic power of LexAgentHallu for evaluating agentic hallucination in law.
Problem

Research questions and friction points this paper is trying to address.

agentic hallucinations
legal agents
multi-step trajectories
benchmark
diagnostic capability
Innovation

Methods, ideas, or system contributions that make the work stand out.

LexAgentHallu
agentic hallucinations
multi-step trajectories
dual-layer taxonomy
fine-grained metrics
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yujin Zhou
Hong Kong University of Science and Technology
M
Mingxuan Zheng
Hong Kong University of Science and Technology
Chuxue Cao
Chuxue Cao
Hong Kong University of Science and Technology
H
Huang Yidan
Hong Kong University of Science and Technology
Jiale Chen
Jiale Chen
Institute of Science and Technology Austria (ISTA)
computer science
Y
Yike Guo
Hong Kong University of Science and Technology
Sirui Han
Sirui Han
The Hong Kong University of Science and Technology
Large Language ModelInterdisciplinary Artificial Intelligence