Rubric-Aligned Disentangled Evaluation of Human Simultaneous Interpreting

📅 2026-09-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究解决了自动评估同声传译的问题,通过构建专业注解语料库并使用改进的COMET-KIWI编码器,提高了与人工评分的相关性。
📝 Abstract
Human simultaneous interpreting (SI) is commonly assessed with analytic rubrics separating meaning transfer, delivery quality, and temporal synchrony, yet no automatic metric is designed for rubric-aligned segment-level SI evaluation. We construct a professionally annotated corpus of 1,101 SI segments with scores for meaning transfer (LQ), delivery quality (EXP), and perceived latency (LAT). We show that structured LLM prompting and scalar supervision collapse rubric dimensions, yielding near-zero correlation with human ratings and strong cross-dimension coupling. To isolate supervision structure under identical backbone capacity, we introduce dual regression heads on a LoRA-adapted COMET-KIWI encoder. On a held-out talk-level test set, the model achieves Pearson correlations of 0.388 (LQ) and 0.301 (EXP), improving over frozen COMET-KIWI. Given low absolute rater agreement, we interpret results relative to human consistency and target stable ranking signals for formative assessment.
Problem

Research questions and friction points this paper is trying to address.

Human Simultaneous Interpreting
Rubric-Aligned Evaluation
Segment-Level
Innovation

Methods, ideas, or system contributions that make the work stand out.

rubric-aligned evaluation
dual regression heads
LoRA-adapted COMET-KIWI
segment-level SI evaluation
💼 Related Jobs
No related jobs found.
Z
Ziyu Zhang
School of Data Science, The Chinese University of Hong Kong, Shenzhen, China
Satoshi Nakamura
Satoshi Nakamura
The Chinese University of Hong Kong, Shenzhen
speech and natural language processing