Building Legal Reward Models for Grounding and Abstention

📅 2026-09-13
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过构建LegalRewardBench解决法律问答中基于证据的推理和在证据不足时的拒绝问题,使用上下文偏好数据改进模型性能。
📝 Abstract
Large language models are increasingly used in high-stakes domains such as law, where systems must ground their reasoning in retrieved evidence and abstain when that evidence is insufficient. However, existing reward models are largely optimised for general preferences rather than contextual grounding, limiting their ability to evaluate these behaviours in retrieval-augmented generation (RAG) settings. We introduce a framework for transforming existing legal QA datasets into contextual preference data and use it to construct LegalRewardBench (LRB), a benchmark for evaluating grounded legal generation under noisy and insufficient retrieval conditions. Across general and legal contextual evaluation, we find that contextual DPO improves grounded evaluation, but performance is sensitive to preference-data construction. Length-balanced augmentation substantially improves grounded legal evaluation, with the strongest configuration combining length-balanced legal and general contextual preference data and improving performance by up to $\mathbf{+25.6}$pp over baseline. We further find evidence of cross-jurisdiction transfer: models contextually refined primarily on Victorian criminal-law data improve grounded evaluation on external US legal benchmarks, including a $\mathbf{+16.2}$pp improvement on \textsc{Housing Statute QA}. Together, these results provide a reproducible foundation for constructing and evaluating grounded legal reward models in retrieval-augmented settings.
Problem

Research questions and friction points this paper is trying to address.

legal QA
contextual grounding
retrieval-augmented generation
reward models
Innovation

Methods, ideas, or system contributions that make the work stand out.

contextual preference data
LegalRewardBench
length-balanced augmentation
cross-jurisdiction transfer
R
Rilton Franzone
University of Oxford
V
Valentin Noël
Devoteam
P
Puyu Wang
University of Oxford
Philip Torr
Philip Torr
Professor, University of Oxford
Department of Engineering
F
Fabio J. Fehr
University of Oxford