Spatial Message Passing in Language Space for Pathology Image Interpretation

📅 2026-08-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of restricted context and fragmented spatial structure in multimodal large language models processing whole slide images (WSIs) by proposing the SLMP framework. Modeling WSIs as spatial text graphs, this approach introduces a novel language-space message passing mechanism that leverages text gradients to optimize interpretable prompts for neighborhood information aggregation. Consequently, it enables adaptive semantic reasoning from cellular to tissue levels without weight fine-tuning. Experiments demonstrate that SLMP improves tumor description accuracy by 3.3 to 19.6 percentage points, significantly enhancing general-purpose model performance while narrowing the gap with specialized models. Ultimately, this work advances both the accuracy and transparency of pathological interpretation within foundation models.
📝 Abstract
Multimodal Large Language Models (MLLMs) can generate pathological descriptions from histological images, but gigapixel Whole Slide Images (WSIs) exceed their visual context limits. The standard tiling workaround makes WSIs tractable yet severs the tissue neighborhoods that define tumor-stroma interfaces and morphology. We introduce Spatial Language Message Passing (SLMP), a framework that performs spatial reasoning entirely in language space, human-readable by construction. SLMP represents a WSI region as a spatial text graph: tiles are nodes initialized with MLLM descriptions, and edges encode spatial adjacency. For each tile, an LLM refines its description by integrating language messages from adjacent tiles under a shared aggregation policy that, on the tile grid, acts as an adaptive local kernel operating on text rather than learned embeddings. This policy is an inspectable prompt that can be refined from model-observed tissue phenotypes via textual gradients, enabling automatic semantic optimization from local cellular context to broader tissue morphology without fine-tuning MLLM weights. On representative HER2 and CAMELYON16 regions, SLMP improves tile-level tumor description accuracy in settings spanning general-purpose and pathology-specialized backbones, with gains of +3.3 to +19.6 percentage points. Random-neighbor ablations confirm that these gains stem from spatial context rather than additional text alone, and inspecting the optimized policies reveals interpretable, tissue-specific decision rules. Besides, without any weight updates or fine-tuning the backbone MLLM, SLMP substantially improves general-purpose MLLMs and narrows its gap to pathology-specialized counterparts, offering a transparent and flexible mechanism for incorporating spatial reasoning into MLLM-based pathology analysis.
Problem

Research questions and friction points this paper is trying to address.

Whole Slide Images
Multimodal Large Language Models
Spatial Reasoning
Pathology Image Interpretation
Visual Context Limits
Innovation

Methods, ideas, or system contributions that make the work stand out.

Spatial Language Message Passing
Textual Gradients
Language Space Reasoning
Spatial Text Graph
Parameter-Efficient Adaptation
J
Jing-Cheng Yang
National Taiwan University, Taipei, Taiwan
H
Hao-Jung Wang
National Taiwan University, Taipei, Taiwan
J
Jinhao Du
University of Oxford, Oxford, United Kingdom
Yang Hu
Yang Hu
Nuffield Department of Medicine, University of Oxford
Deep LearningKnowledge GraphComputational Pathology
M
Ming-shan Tsai
ZYTCA Limited, Oxford, United Kingdom
J
Jens Rittscher
University of Oxford, Oxford, United Kingdom
B
Bin Li
University of Oxford, Oxford, United Kingdom