Comparison of Unsupervised Metrics for Evaluating Judicial Decision Extraction

📅 2025-10-02
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Legal NLP lacks human-annotated data for evaluating semantic block extraction quality in judicial judgments. Method: We systematically benchmark 16 unsupervised metrics across seven legal content extraction tasks, integrating document-level, semantic-level, structural-level, pseudo-labeling, and law-specific measures—including TF-Coherence, Coverage Ratio, Legal Term Density, and LLM-based scoring—validated via bootstrap correlation analysis, Lin’s concordance correlation coefficient, and mean absolute error (MAE) against expert judgments. Results: TF-Coherence and Coverage Ratio achieve the strongest alignment with expert ratings (r > 0.5, MAE < 0.14), significantly outperforming LLM scoring (r = 0.382), thereby exposing LLMs’ limitations in fine-grained legal assessment. This work introduces the first scalable, unsupervised evaluation framework tailored to legal text, balancing computational efficiency and reliability, and enabling annotation-free, automated quality screening for large-scale legal NLP systems.

Technology Category

Application Category

📝 Abstract
The rapid advancement of artificial intelligence in legal natural language processing demands scalable methods for evaluating text extraction from judicial decisions. This study evaluates 16 unsupervised metrics, including novel formulations, to assess the quality of extracting seven semantic blocks from 1,000 anonymized Russian judicial decisions, validated against 7,168 expert reviews on a 1--5 Likert scale. These metrics, spanning document-based, semantic, structural, pseudo-ground truth, and legal-specific categories, operate without pre-annotated ground truth. Bootstrapped correlations, Lin's concordance correlation coefficient (CCC), and mean absolute error (MAE) reveal that Term Frequency Coherence (Pearson $r = 0.540$, Lin CCC = 0.512, MAE = 0.127) and Coverage Ratio/Block Completeness (Pearson $r = 0.513$, Lin CCC = 0.443, MAE = 0.139) best align with expert ratings, while Legal Term Density (Pearson $r = -0.479$, Lin CCC = -0.079, MAE = 0.394) show strong negative correlations. The LLM Evaluation Score (mean = 0.849, Pearson $r = 0.382$, Lin CCC = 0.325, MAE = 0.197) showed moderate alignment, but its performance, using gpt-4.1-mini via g4f, suggests limited specialization for legal textse. These findings highlight that unsupervised metrics, including LLM-based approaches, enable scalable screening but, with moderate correlations and low CCC values, cannot fully replace human judgment in high-stakes legal contexts. This work advances legal NLP by providing annotation-free evaluation tools, with implications for judicial analytics and ethical AI deployment.
Problem

Research questions and friction points this paper is trying to address.

Evaluating unsupervised metrics for legal text extraction quality
Assessing extraction of semantic blocks from judicial decisions
Comparing metric performance against expert human evaluations
Innovation

Methods, ideas, or system contributions that make the work stand out.

Evaluated unsupervised metrics for legal text extraction
Used novel formulations to assess semantic block quality
Applied bootstrapped correlations and concordance coefficients
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
I
Ivan Leonidovich Litvak
Dubna, Russian Federation.
A
Anton Kostin
Moscow Center for Advanced Studies, Russian Federation.
F
Fedor Lashkin
Novorossiysk, Russian Federation.
T
Tatiana Maksiyan
Moscow, Russian Federation.
S
Sergey Lagutin
Moscow, Russian Federation.