Abstract
Learning transferable region embeddings is fundamental to urban computing, supporting applications from economic forecasting to public safety. Existing multi-view fusion methods predominantly rely on coarse-grained fusion, applying uniform weights across an entire city or task, thereby neglecting spatial heterogeneity—the fact that the dominant view driving a region’s function varies significantly across regions. To address this, we propose RAMEN, a two-stage framework for Region-adaptive Mixture of Ego-Networks learning. In the pre-training stage, a unified spatially aware Transformer distills universal urban semantics. In the fine-grained adaptation stage, a Mixture of Ego-Networks (MoEN) module employs a region-adaptive gating mechanism to dynamically allocate exclusive view weights for each region. A Region Ego-aware Spatial Transformer (REST) then aggregates these fused local subgraphs by explicitly injecting degree and physical distance priors, overcoming the over-smoothing limitations of traditional GNNs. Extensive experiments on real-world datasets for check-in, crime, service call and population prediction show that RAMEN consistently outperforms state-of-the-art baselines, achieving up to 35.3% MAE improvement. Visualizations of gating weights further suggest that RAMEN’s fusion aligns well with real-world urban physical characteristics.
IPC Classification
Keywords
€ 4.00