MARS-CLIP: Multi-Resolution and Attention Refined Zero-Shot Image Segmentation

📅 2026-09-01
🏛️ International Conference on Information Photonics
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决CLIP在零样本图像分割任务中因分辨率低和结构信息丢失导致的问题,提出MARS-CLIP框架,通过多分辨率特征提取和注意力细化机制提升分割精度。
📝 Abstract
Contrastive Language-Image Pre-training (CLIP) has demonstrated impressive capabilities in zero-shot transfer but often struggles with dense prediction tasks due to low spatial resolution and the loss of structural information. To address these limitations, we propose MARS-CLIP (Multi-resolution and Attention Refined Segmentation for CLIP), a novel framework for zero-shot semantic segmentation. Our approach introduces two key strategies: (i) a multi-resolution feature extraction module that fuses local fine-grained features with global context to overcome input resolution constraints, and (ii) an attention refinement mechanism that injects spatial and color biases from intermediate layers into the final self-attention block to accurately restore object boundaries. A set of experiments on six public datasets demonstrates that MARS-CLIP significantly outperforms state-of-the-art methods.
Problem

Research questions and friction points this paper is trying to address.

Zero-Shot Image Segmentation
Low Spatial Resolution
Structural Information Loss
Innovation

Methods, ideas, or system contributions that make the work stand out.

multi-resolution feature extraction
attention refinement mechanism
zero-shot semantic segmentation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
N
Nagito Saito
Graduate School of Information Sciences, Tohoku University, Japan
S
Shintaro Ito
Graduate School of Information Sciences, Tohoku University, Japan
Koichi Ito
Koichi Ito
Associate Professor, Graduate School of Information Sciences, Tohoku University
Image ProcessingComputer VisionBiometrics
T
Takafumi Aoki
Graduate School of Information Sciences, Tohoku University, Japan