JPO: Juris Policy Optimization for Structured Legal Reasoning in Criminal Judgment Prediction

📅 2026-08-30
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文针对刑事判决预测中的结构化法律推理问题,提出了一种基于强化学习的方法Juris Policy Optimization (JPO),通过优化法律预测质量、推理结构完整性和跨步骤一致性来改进模型性能。
📝 Abstract
Criminal judgment prediction requires models to infer statutory articles, charges, and sentencing outcomes from case facts. Unlike standard classification tasks, it involves a structured reasoning process in which statutes should be matched with facts, charges should be justified by statutes, and sentencing outcomes should remain consistent with charges. Existing approaches optimize final labels, and while some have attempted to evaluate reasoning quality, their evaluations are indirect, often relying on LLM-generated rubrics that reflect model-internal preferences rather than the inherent logical structure of legal adjudication. We propose Juris Policy Optimization (JPO), a post-training framework for structured legal reasoning in Chinese criminal judgment prediction. JPO first uses teacher-generated rationales to supervise a standardized four-step reasoning process, and then applies reinforcement learning with a composite reward over legal prediction quality, reasoning structure completeness, and cross-step consistency. JPO further introduces token-level advantage reweighting and adaptive clipping for legally salient reasoning segments. Experiments on multiple open-source language models and three Chinese legal benchmarks show that JPO consistently improves both judgment prediction and reasoning quality over supervised fine-tuning and reinforcement learning baselines.
Problem

Research questions and friction points this paper is trying to address.

criminal judgment prediction
structured reasoning
legal adjudication
reasoning quality
reinforcement learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Juris Policy Optimization
structured legal reasoning
reinforcement learning
composite reward
adaptive clipping
Z
Zhaolu Kang
Tencent
Yantao Liu
Yantao Liu
Qwen, Alibaba
Reinforcement LearningReward ModelingLarge Language Models
T
Tailong Luo
Peking University
L
Leqi Zheng
Tsinghua University
L
Lei Wei
Peking University
C
Chenghua Zhu
Peking University
J
Junhao Gong
Peking University
J
Jiachen Qian
City University of Hong Kong
E
Eric Hanchen Jiang
University of California
J
Jiaxin Liu
University of Illinois Urbana-Champaign
Yuan Wang
Yuan Wang
Zhejiang University
Medical MLLM
H
Hao Zhang
Peking University
Z
Zixia Wang
Peking University
R
Rong Fu
Peking University
Z
Zheng Lin
University of Hong Kong
R
Richeng Xuan
Tencent
Z
Zhichao Hu
Tencent