Model Effect or Label Effect? Refined Annotations and a Human-Referenced Benchmark for Pulmonary Embolism Segmentation

📅 2026-08-25
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过改进标注和建立人类参考框架,量化了评估标注对肺栓塞分割性能的影响,并使用nnU-Net模型比较了标注效果和模型训练效果。
📝 Abstract
Purpose: To quantify how evaluation annotations influence measured pulmonary embolism (PE) segmentation performance relative to model training changes, and to establish a human-referenced framework. Materials and Methods: This retrospective study screened 166 voxel-annotated CT pulmonary angiography cases from CADPE (n=91), FUMPE (n=35), and READ (n=40); 149 were included. A primary rater annotated PE by protocol, and a senior thoracic radiologist reviewed and revised all segmentations. Three additional raters at three centers annotated a 15-case subset. The label effect was measured by evaluating two pretrained nnU-Net models (nnU-Net-A, nnU-Net-B) against original and refined annotations. The model effect was measured by comparing the same architecture trained on different dataset combinations with annotations fixed. The benchmark model (nnPE) was trained with leave-one-dataset-out and pooled five-fold cross-validation. Four metric categories were analyzed with case-paired Wilcoxon signed-rank tests, Benjamini-Hochberg correction, and bootstrap 95% CIs. Results: Changing only the annotation increased mean DSC by 0.143 (0.122-0.166) for nnU-Net-A and 0.188 (0.163-0.213) for nnU-Net-B (both P < .001), whereas changing training-dataset composition changed DSC by 0.028. The label effect exceeded the model effect on CADPE and FUMPE and was 0.045 on READ. Within-mask attenuation SD fell in all three datasets after re-annotation (all P < .001). nnPE reached DSC 0.72 +/- 0.22 on pooled cross-validation but scored below all four annotators across 52 paired comparisons (all corrected P < .05). Conclusion: Evaluation annotations affected measured PE segmentation performance at least as much as model training choices. A human-referenced evaluation framework is publicly available for future study.
Problem

Research questions and friction points this paper is trying to address.

Pulmonary Embolism
Segmentation
Annotation
Model Training
Benchmark
Innovation

Methods, ideas, or system contributions that make the work stand out.

refined annotations
human-referenced benchmark
pulmonary embolism segmentation
label effect
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Q
Qihang Sun
Technical University of Munich (TUM), Munich, Germany; Munich Center for Machine Learning (MCML), Munich, Germany
Z
Zhongxiao Liu
Department of Radiology, The Affiliated Hospital of Xuzhou Medical University, Xuzhou, China
Bailiang Jian
Bailiang Jian
Techinical Unversity of Munich
medical image registration
S
Shenman Qiu
Department of Radiology, The Affiliated Hospital of Xuzhou Medical University, Xuzhou, China
J
Jingyuan Wang
Department of Radiology, The Third Affiliated Hospital of Soochow University, Changzhou, China
L
Lei Zhang
Department of Radiology, The Affiliated Taizhou People’s Hospital of Nanjing Medical University, Taizhou, China
L
Lixiang Xie
Department of Radiology, The Affiliated Hospital of Xuzhou Medical University, Xuzhou, China
Jiazhen Pan
Jiazhen Pan
Technical University of Munich
Machine LearningMedical Image ComputingBiomedical Image Analysis
Christian Wachinger
Christian Wachinger
Technical University of Munich
AI in Medical ImagingGeometric Deep LearningCausal InferenceMulti-Modal Diagnostics