Is this Citation on Point?

📅 2026-08-12
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of determining whether legal citations genuinely support the claims they are intended to substantiate. The authors propose an evaluation framework based on controlled perturbations—specifically, substituting cited cases or altering page numbers—to systematically assess models’ ability to distinguish between topical relevance and proposition-level evidentiary support. Experiments were conducted across two legal corpora (court opinions and legal briefs) using fourteen model configurations enhanced with high-reasoning prompting. Results reveal that while models excel at identifying irrelevant cases (93–100% accuracy), they exhibit marked deficiencies in detecting page-number mismatches (37–83% accuracy). Notably, even GPT-5.4 in high-reasoning mode fails to flag 40% of such page-level errors, underscoring a critical limitation: current models conflate thematic relatedness with precise, citation-specific support.
📝 Abstract
In 2023, a New York judge sanctioned two attorneys in Mata v. Avianca for filing a brief with hallucinated citations generated by ChatGPT. Such failures are largely caught by database lookups; the harder problem is detecting citations that point to real cases but do not support the propositions for which they are offered -- a failure mode that existing evaluations of LLMs for legal use cases largely overlook. In this paper, we study proposition-level citation support verification through controlled perturbations of real legal citations obtained from two legal corpora, either replacing the cited case or changing only the pinpoint page within the same case. We evaluate fourteen model configurations on the resulting examples. Models catch 93-100% of wrong-case corruptions. They catch only 37-61% of wrong-pinpoint corruptions on court opinions and 52-83% on legal briefs. When models fail to catch wrong-pinpoint corruptions, they accept the citation based on topical overlap rather than page-level support. Scale and extended reasoning narrow the gap but do not close it: GPT-5.4 with high reasoning effort still misses 40% of pinpoint mismatches on court opinions and 18% on briefs. Prompting the model to verify support at the cited page improves recall, but it also raises the false positive rate. Recognizing the right legal topic and verifying support for the cited proposition are distinct capabilities, and current models conflate them.
Problem

Research questions and friction points this paper is trying to address.

citation support verification
legal reasoning
hallucinated citations
pinpoint citation
proposition-level verification
Innovation

Methods, ideas, or system contributions that make the work stand out.

citation support verification
legal hallucination
pinpoint citation
controlled perturbation
large language models
🔎 Similar Papers
No similar papers found.
A
Apurv Verma
Bloomberg, New York, NY, USA