Learning Where to Focus: Self-Supervised Multi-Scale ViTs for Histopathology

📅 2026-09-16
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过提出CRAFT框架,利用自监督学习在病理学图像中动态分配空间分辨率,以提高病理表示的质量。
📝 Abstract
Pathologists diagnose diseases by first locating suspicious tissue and then examining it at higher magnification, whereas self-supervised vision transformers (ViTs) allocate the same spatial resolution to every image region despite diagnostic evidence being sparse and spanning multiple biological scales. Recent pathology foundation models have substantially improved representation quality by scaling training data and model capacity, but largely retain uniform tokenization. We instead investigate whether pathology representations can be improved by learning where to allocate spatial resolution during self-supervised learning. To this end, we propose CRAFT (Coarse-to-fine Region-Adaptive Feature Tokenization), a DINO-based framework that learns image-dependent mixed-scale representations by using self-supervised attention to selectively refine informative regions while preserving coarse context, together with a symmetric cross-scale regularization objective that encourages complementary coarse and fine representations. Across CAMELYON16, TCGA-Lung subtype classification, and TCGA-LUAD survival prediction, CRAFT consistently outperforms comparable-scale self-supervised methods while requiring lower inference computation. Despite using only a compact 22M parameter backbone trained on comparatively small pathology datasets, CRAFT remains competitive with, and often surpasses, substantially larger pathology foundation models.
Problem

Research questions and friction points this paper is trying to address.

self-supervised vision transformers
histopathology
spatial resolution
diagnostic evidence
multi-scale
Innovation

Methods, ideas, or system contributions that make the work stand out.

CRAFT
self-supervised attention
mixed-scale representations
cross-scale regularization
🔎 Similar Papers
No similar papers found.