SlideMix: Enhancing Whole Slide Image Analysis via Multimodal Shuffling

📅 2026-08-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
针对WSI分析中的挑战,提出SlideMix框架,通过多模态混洗增强MIL方法,提高诊断相关区域识别和跨尺度特征学习的准确性。
📝 Abstract
Histopathological whole slide images (WSIs) are central to cancer diagnosis, but their gigapixel scale, tissue heterogeneity, weak slide-level supervision, sparse diagnostic regions, and multi-scale evidence make robust automated analysis challenging. Multiple instance learning (MIL) is widely used to aggregate tile-level features into slide-level predictions, yet existing augmentation strategies often perturb tissue regions without preserving diagnostic relevance, slide context, or cross-scale structure. We propose SlideMix, a model-agnostic multimodal augmentation framework for MIL-based WSI analysis. SlideMix uses a retrieval-augmented vision-language model (VLM)-based Visual-Language Adaptive Region selector to identify diagnostically relevant regions and reduce weak-label noise. It then performs In-place Tile Shuffling within meaningful tissue regions to mix feature embeddings while preserving slide-level context. A VLM-based soft-labeling module supervises mixed samples, while a multi-factor, loss-driven online Curriculum-Learning Feedback scheme adaptively controls shuffle granularity, feature similarity, and shuffle ratio to promote cross-scale representation learning. Across 11 WSI datasets comprising 20,523 slides, 8 diagnostic tasks, and 10 WSI backbones, SlideMix improves accuracy and generalization in most settings and compares favorably with established augmentation baselines, providing a simple plug-and-play approach for more robust and scalable digital pathology models. Source code: https://github.com/Xia-Research-Lab/SlideMix
Problem

Research questions and friction points this paper is trying to address.

Whole Slide Images
Multimodal Shuffling
Multiple Instance Learning
Tissue Heterogeneity
Weak Supervision
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multimodal Augmentation
Visual-Language Model
In-place Tile Shuffling
Curriculum-Learning Feedback
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
C
Chad Wong
Department of Electrical Engineering and Computer Science, University of California, Irvine
S
Sicheng Chen
Department of Electrical Engineering and Computer Science, University of California, Irvine
Tianyi Zhang
Tianyi Zhang
Assistant Professor of Computer Science, Purdue University
Software EngineeringHuman-Computer InteractionLarge Language Models
E
Enhui Chai
PuzzleLogic Pte Ltd, Singapore 229594, Singapore
Yueming Jin
Yueming Jin
Assistant Professor, National University of Singapore
Medical Image AnalysisSurgical AI&RoboticsMultimodal Learning
Z
Zeyu Liu
PuzzleLogic Pte Ltd, Singapore 229594, Singapore
Fei Xia
Fei Xia
Assistant Professor, University of California, Irvine
Optical ImagingComputational ImagingNeuromorphic ComputingNeurophotonics