Benchmarking Vision-Language Models for Automated Pathology Diagnosis and Report Generation

📅 2026-09-01
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决WSI病理报告生成数据稀缺和复杂性问题,构建了泛亚洲WSI-报告数据集,并通过MICCAI挑战赛评估多种模型方法。
📝 Abstract
The rapid advancement of vision-language models (VLMs) has accelerated progress in computational pathology; however, whole-slide image (WSI)-based pathology report generation remains limited by the scarcity of large-scale WSI--report datasets and the complexity of mapping spatially distributed visual patterns to structured clinical text. To address this, we introduce a clinically curated Pan-Asia WSI--report dataset of approximately 10,500 pairs from five institutions and establish the REG 2025 benchmark through a MICCAI challenge for systematic evaluation of multimodal models. We analyze submitted methods spanning pretrained VLMs, multiple-instance learning frameworks, hierarchical expert models, retrieval-augmented generation, and cross-modal Transformers. Rather than indicating that VLM use alone was sufficient for superior performance, the results suggest that top-performing methods benefited from structured report representations, hierarchical diagnostic decomposition, and effective multimodal grounding. We identify key limitations, including instability in quantitative attribute estimation (e.g., numeric hallucination) and a tendency toward diagnostic overspecification, with some errors resembling known diagnostic pitfalls in routine pathology. These findings establish REG 2025 as a benchmark for evaluating WSI-based structured report generation and vision-language understanding in computational pathology, providing insights for the design of clinically grounded multimodal pathology models.
Problem

Research questions and friction points this paper is trying to address.

vision-language models
whole-slide image
pathology report generation
multimodal models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language Models
Whole-Slide Image
Structured Report Representations
Hierarchical Diagnostic Decomposition
Multimodal Grounding
💼 Related Jobs
No related jobs found.
Y
Yumi Lee
H
Harim Oh
H
Hyoryung Kim
M
Minji Kim
Eunsu Kim
Eunsu Kim
KAIST
AINLP
H
Hyeseong Lee
J
Junya Fukuoka
A
Andrey Bychkov
J
Jijgee Munkhdelger
R
Rajiv Kumar Kaushal
A
Ayushi Sahay
R
Rajni Yadav
B
Bharathi Prabakaran
S
Sulen Sarioglu
S
Serdar Balcı
I
Ilknur Turkmen
Yuri Tolkach
Yuri Tolkach
Institute of Pathology, University Clinic of Cologne
Artificial IntelligenceDigital PathologyOncologyGU pathology
C
Christian Harder
J
Julian Westerdorf
R
Reinhard Buettner
A
Audun Ljone Henriksen
S
Sepp De Raedt
B
Byung Hyun Lee
S
Sungjin Lim
J
Joohoon Lee