Hierarchical Adaptive Feature Refinement Network for VHR Remote Sensing Image Segmentation

πŸ“… 2026-08-16
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the challenges of ineffective multi-stage feature utilization and insufficient structural guidance in semantic segmentation of high-resolution remote sensing imagery by proposing HAFR-Net, a progressive refinement framework. The method introduces a heterogeneous guided fusion mechanism to adaptively optimize hierarchical features and incorporates a frequency-domain residual adapter to enhance detail representation. Furthermore, a confusion-aware tri-prior decoder is constructed to strengthen boundary constraints. Extensive evaluations on four benchmark datasets demonstrate that HAFR-Net achieves significant improvements in mIoU and outperforms existing baselines in boundary preservation and fine-structure segmentation accuracy under complex scenarios. These results confirm the framework’s effectiveness in overcoming critical bottlenecks in remote sensing image segmentation.
πŸ“ Abstract
Semantic segmentation of very-high-resolution (VHR) remote sensing imagery increasingly benefits from strong pretrained hierarchical encoders, yet exploiting their multi-stage representations remains difficult. Nearby regions demand different balances between fine detail and semantic context, aggressive task-specific transformations perturb useful pretrained features, and conventional semantic supervision provides limited structural guidance. We present HAFR-Net, a progressive refinement framework that adaptively organizes and conservatively refines hierarchical representations instead of replacing them with a monolithic decoder transformation. Heterogeneity-Guided Stage-Adaptive Fusion (HG-SAF) predicts dense stage weights conditioned on local feature variation. A Frequency-Residual Adapter (FRA) then injects frequency information through a bounded, zero-initialized residual branch that keeps the fused representation as its reference. A Confusion-Aware Tri-Prior Decoder (CATP) finally regularizes the prediction with boundary, objectness, and training-derived class-relation cues. Under a matched Swin-B training and single-scale inference protocol, HAFR-Net attains 84.12%, 87.86%, 55.17%, and 67.70% mIoU on ISPRS Vaihingen, ISPRS Potsdam, LoveDA, and OpenEarthMap, improving the matched UPerNet baseline by 0.55, 0.95, 1.55, and 1.84 percentage points, respectively. Controlled analyses further show consistent spatial reweighting beyond content-only routing, improved boundary and thin-structure accuracy over matched spatial and spectral alternatives, and reduced confusion on pre-declared class pairs.
Problem

Research questions and friction points this paper is trying to address.

VHR remote sensing image segmentation
hierarchical representations
pretrained encoders
semantic supervision
Innovation

Methods, ideas, or system contributions that make the work stand out.

Hierarchical Adaptive Feature Refinement
Heterogeneity-Guided Stage-Adaptive Fusion
Frequency-Residual Adapter
Confusion-Aware Tri-Prior Decoder
VHR Remote Sensing Segmentation
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
S
Shuaishuai Cao
Meng Tang
Meng Tang
Assistant Professor, University of California, Merced
computer visionmachine learningoptimization
S
Shuwei Peng
X
Xuan Liu
M
Min Huang
J
Jie Chen
J
Jiacheng Niu
Y
Yong Chen
E
Edore Akpokodje
H
Hui Lin