🤖 AI Summary
This study addresses the limitation of uneven street view coverage in large-scale urban perception measurement by proposing CVLNet. Integrating AlphaEarth embeddings with multi-source urban contextual data, the method employs cross-view learning and an adaptive gating mechanism to enable city-wide subjective perception mapping without relying on street view imagery during inference. Experiments across four cities demonstrate that the model achieves an R² of 0.76, outperforming baselines by 5.9%–11.3%. CVLNet successfully extends perceptual coverage to entire road networks and effectively reveals environmental exposure inequalities. These findings establish CVLNet as an efficient paradigm for scalable urban perception research, overcoming data sparsity constraints inherent in traditional street view-based approaches.
📝 Abstract
Residents' perception of the urban streetscape is an important factor in public health, active mobility, and social wellbeing. Street view imagery (SVI) has emerged as a widely used data source for assessing these perceptual qualities, yet its uneven coverage and irregular updating limit large-scale measurement. Here, we present CVLNet, a Cross-View Learning Network that predicts street-level perception from AlphaEarth embeddings and multi-source urban contextual data without requiring SVI at inference. CVLNet applies per-task adaptive gating to jointly model five perceptual dimensions, using labels from the pretrained SVI-Percept model as ground truth. The proposed method is evaluated across four Southeast Asian cities: Singapore, Kuala Lumpur, Jakarta, and Manila. CVLNet achieves a median road-segment-level Adjusted $R^{2}$ of 0.76 and consistently outperforms the baseline models, with gains ranging from 5.9--11.3% across the five perceptual dimensions. Ablation experiments show that AlphaEarth features and urban contextual features contribute complementary information. We further produce citywide road-level streetscape perception maps for five subjective perceptual dimensions across all four cities, extending perception estimation from the 13--31% of the road network directly covered by available SVI to the complete road network of each city. Integrating these maps with WorldPop gridded population data, we quantify exposure inequality across population-density, demographic, and land-use groups using the Deficit Palma Ratio. These results demonstrate that remote sensing can serve as a scalable alternative to SVI for citywide streetscape perception mapping, enabling a more comprehensive assessment of urban environmental inequality.