Improving Argument Saliency Coverage in Small LLMs for Long Legal Opinion Summarization via Sequence-Level Distillation

📅 2026-08-30
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
通过序列级蒸馏方法,从长上下文教师模型中提取知识,以提高小型语言模型在法律意见摘要中的论点显著性覆盖率,此方法数据效率高且无需额外标注。
📝 Abstract
We show that sequence-level distillation from a capable long-context teacher model is a simple, annotation-free, and data-efficient strategy for improving argument saliency coverage in long legal opinion summarization, where small LLMs often struggle to retain the most salient argumentative content. Across student model sizes, distillation consistently surpasses tuning on expert-written summaries in our legal-opinion setting. We further demonstrate that most gains are achieved with as few as ~10 training summaries, highlighting the strong data efficiency of teacher-generated supervision. Finally, we find that summary distillation is sufficient for improvements: reasoning-chain distillation remains competitive with summary-only distillation, but provides marginal benefit when combined with summary supervision.
Problem

Research questions and friction points this paper is trying to address.

Argument Saliency Coverage
Small LLMs
Long Legal Opinion Summarization
Innovation

Methods, ideas, or system contributions that make the work stand out.

sequence-level distillation
long-context teacher model
argument saliency coverage
data efficiency
legal opinion summarization