Enhancing Pathological VLMs with Cross-scale Reasoning

📅 2026-06-15
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing pathological vision-language models lack an explicit cross-scale reasoning mechanism, hindering their ability to integrate multi-scale evidence—from tissue to cellular levels—for precise diagnosis. This work introduces the first training and evaluation framework tailored for multi-magnification reasoning in histopathology. It proposes a leakage-aware data construction pipeline, designs adversarial text filtering and constraint-guided question generation strategies, and establishes Scale-VQA, a high-quality multi-scale visual question answering benchmark. Building upon this, the authors develop ScaleReasoner-R1, a model trained via reinforcement learning. ScaleReasoner-R1 achieves state-of-the-art performance on the newly curated cross-scale benchmark and outperforms existing methods on multiple single-scale pathological VQA tasks, demonstrating that even limited cross-scale supervision can substantially enhance pathological understanding.
📝 Abstract
Pathological images are inherently multi-scale, requiring pathologists to integrate evidence from global tissue architecture at low magnification to cellular morphology at higher magnification for accurate diagnosis. While existing pathological datasets for vision-language model (VLM) include various scales, they often lack an explicit cross-scale reasoning objective. This limitation prevents VLMs from capturing essential cross-scale representations and learning evidence-based reasoning. To bridge this gap, we introduce the first cross-scale training and evaluation paradigm that formulates pathology interpretation as multi-magnification reasoning. However, creating such a task reveals a critical challenge: multi-image visual question answering (VQA) is prone to text-only shortcuts, which allow models to guess answers using magnification-dependent artifacts rather than visual evidence. To address this, we propose a leakage-aware curation pipeline that combines adversarial text-only screening with constraint-guided question design. Using this pipeline, we construct Scale-VQA, a high-quality benchmark with 4,685 multiple-choice questions grounded in 2,537 pathology images across multiple magnification levels. Finally, we present ScaleReasoner-R1, a model trained via reinforcement learning to optimize performance on the cross-scale VQA task. ScaleReasoner-R1 achieves state-of-the-art performance on our cross-scale reasoning benchmark and generalizes to SOTA performance on established single-scale benchmarks. Findings suggest that even the limited cross-scale supervision can significantly improve pathological understanding. The code and demos will be open-sourced.
Problem

Research questions and friction points this paper is trying to address.

cross-scale reasoning
pathological VLMs
multi-magnification
visual question answering
text-only shortcuts
Innovation

Methods, ideas, or system contributions that make the work stand out.

cross-scale reasoning
pathological VLM
multi-magnification VQA
leakage-aware curation
reinforcement learning
Chi Phan
Chi Phan
Associate Professor, Curtin University
Chemical EngineeringInterface ScienceSurfactantEmulsion Surface Charge
Tianyi Zhang
Tianyi Zhang
PhD student at NUS, Singapore
Medical Image AnalysisDeep LearningComputer Vision
Q
Qiaochu Xue
Department of Electrical and Computer Engineering, National University of Singapore, Singapore
Y
Yufeng Wu
PuzzleLogic Pte Ltd, Singapore
D
Dan Hu
Department of Pathology, Fujian Medical University Cancer Hospital & Fujian Cancer Hospital, Fuzhou, China
Z
Zeyu Liu
PuzzleLogic Pte Ltd, Singapore
S
Sudong Wang
PuzzleLogic Pte Ltd, Singapore
Yueming Jin
Yueming Jin
Assistant Professor, National University of Singapore
Medical Image AnalysisSurgical AI&RoboticsMultimodal Learning