EVEREST:Endogenous Vision-Language Reinforcement Reasoning Exploration for Urban Socio-Semantic Segmentation

📅 2026-08-25
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决城市社会语义分割中目标边界不准确的问题,提出EVEREST模型,采用自我中心探索策略并结合强化学习以主动探索和修正边界线索。
📝 Abstract
Urban socio-semantic segmentation leverages digital and satellite imagery to provide critical spatial semantic information for downstream applications such as urban resource allocation. Although existing methods achieve high segmentation accuracy, they still suffer from inaccurate delineation of target boundaries. The underlying issue is that current models primarily rely on passively aggregated global cross-modal cues, lacking active exploration of the environment. To address this limitation, we propose the EVEREST model, which adopts an egocentric exploration strategy that enables the model to actively investigate boundary cues and perform self-correction. In addition, we formulate discrete natural-language prompts as pseudocode to regularize the execution logic. Reinforcement learning is further employed to implement this irreducible process and elicit the model's structured reasoning capability. Our EVEREST achieves optimal performance on all metrics in the real world urban socio-semantic dataset, demonstrating the superiority of our model. Codes are available at https://anonymous.4open.science/r/EVEREST-9D21/.
Problem

Research questions and friction points this paper is trying to address.

Urban Socio-Semantic Segmentation
Boundary Delineation
Active Exploration
Innovation

Methods, ideas, or system contributions that make the work stand out.

Egocentric Exploration
Natural-Language Prompts
Reinforcement Learning
Structured Reasoning