ContractScrub: A benchmark for final review of legal contracts

📅 2026-08-20
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决法律合同最终审查自动化问题,提出ContractScrub基准,评估大语言模型在检测错误和不一致方面的表现。
📝 Abstract
Legal work, with its heavy reliance on processing large amounts of text, is often considered one of the domains most exposed to the use of LLMs. Contract ``scrubbing,'' the final review of transactional agreements for errors and inconsistencies, is a particularly suitable task for automation, because it is routine, painstaking work requiring detailed attention to long documents. Scrubbing also seems to align naturally with the general capabilities expected of frontier LLMs around long-context reasoning, consistency checking, and named entity recognition (NER). Despite the economic value and potential for automation, no formal evaluations of LLMs performing contract scrubbing have been conducted. We introduce ContractScrub, the first benchmark designed to evaluate contract scrubbing capabilities, comprising contracts hand-crafted by experienced lawyers over diverse error categories such as misuse of defined terms, incorrect references, and inconsistent language. Frontier models perform surprisingly poorly with only one model reaching 0.75 macro average recall despite strong performance on seemingly related general benchmarks, demonstrating the practical limits of current models and the importance of narrowly targeted, domain-specific benchmarks for measuring real-world impact.
Problem

Research questions and friction points this paper is trying to address.

Legal Contracts
Contract Scrubbing
Automation
LLMs
Benchmark
Innovation

Methods, ideas, or system contributions that make the work stand out.

ContractScrub
legal contracts
benchmark
LLMs
automation
Yejin Bang
Yejin Bang
Ph.D. Candidate, HKUST
LLM EvaluationNLPResponsible AI
K
Kirsty Fielding
Thomson Reuters Foundational Research, London, UK
B
Brandan Oliver
Thomson Reuters Foundational Research, London, UK
B
Brian Birke
Thomson Reuters Foundational Research, London, UK
Nabeel Seedat
Nabeel Seedat
University of Cambridge
Machine LearningUncertainty QuantificationData-Centric AILarge Language ModelsAI for health
A
Andrew M. Bean
Thomson Reuters Foundational Research, London, UK; Imperial College London, UK