FIRM: Fine-Grained Intra-Token Representation of Masks for Remote Sensing Reasoning Segmentation

📅 2026-08-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the issues of coarse boundaries and instance adhesion in remote sensing semantic segmentation caused by multi-target mixing within visual tokens. To overcome these limitations, we propose FIRM, a novel method that innovatively introduces intra-token sub-unit mask representations and a lightweight continuous rendering mechanism. By transcending single-label constraints through sub-unit prediction, lookup table transformation, and soft structural field marginalization, FIRM achieves fine-grained segmentation. Extensive experiments demonstrate state-of-the-art performance across five benchmarks. Notably, on the LASER dataset, FIRM attains GIoU/CIoU scores of 70.5/80.5 and improves the EarthReason metric by 3.0 points, significantly enhancing segmentation accuracy in complex scenes.
📝 Abstract
Reasoning segmentation requires multimodal large language models (MLLMs) to translate implicit instructions into precise pixel-level masks. MLLMs encode an image as visual tokens, each of which merges a group of image patches. In remote sensing images, small targets, thin structures, and adjacent instances can occupy different parts of the same visual token. Assigning a single binary mask label to such a token loses its internal spatial structure, causing nearby targets to merge and object boundaries to become coarse. To bridge this representational gap, we introduce FIRM, a Fine-grained Intra-token Representation of Masks. For each visual token, FIRM predicts a mask code that specifies an $r\times r$ binary sub-cell pattern rather than a single foreground/background label. Given a target identified by the MLLM, the complete grid of mask codes is predicted in one mask pass. Fixed lookup converts the predicted codes into a discrete sub-cell mask, while marginalizing the code distribution yields a soft structural field. To further recover fine-grained boundaries within each sub-cell, we introduce a lightweight continuous renderer that refines this field using pre-merge visual features and image details. Across five reasoning and referring segmentation benchmarks on satellite and UAV images, FIRM achieves leading results, including $70.5/80.5$ gIoU/cIoU on LaSeRS and a $3.0$-point average gain on EarthReason. These results demonstrate the value of explicitly representing intra-token mask patterns for fine-grained MLLM segmentation.
Problem

Research questions and friction points this paper is trying to address.

Remote Sensing Reasoning Segmentation
Intra-Token Representation
Visual Tokens
Fine-Grained Boundaries
MLLMs
Innovation

Methods, ideas, or system contributions that make the work stand out.

Fine-Grained Intra-Token Representation
Mask Code Prediction
Continuous Renderer
Remote Sensing Reasoning Segmentation
Multimodal Large Language Models
🔎 Similar Papers
2024-09-20IEEE Transactions on Geoscience and Remote SensingCitations: 2
💼 Related Jobs
No related jobs found.
W
Weidong Tang
Xi’an Jiaotong University, Xi’an, China
Kaiyu Li
Kaiyu Li
Wilfrid Laurier University, Canada
Data governance and Data preparationData market and Data economy
Y
Yikai Wang
Renmin University of China, Beijing, China
Yanan Wu
Yanan Wu
China Medical University | NEU (PhD) | CUHK (RA)
Medical Image Analysis
H
Haotian Gan
Shaanxi University of Science and Technology, Xi’an, China
S
Shihong Wang
Xi’an Jiaotong University, Xi’an, China
X
Xiangyong Cao
Xi’an Jiaotong University, Xi’an, China