CityPlanner: A Sandbox Agent for Executable Urban Planning

📅 2026-09-08
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出CityPlanner,通过UrbanSandbox环境和原子任务强化学习方法解决城市规划中的空间优化问题,实验表明优于现有方法。
📝 Abstract
Urban planning is a real-world spatial optimization problem that requires selecting feasible actions from large candidate spaces under practical objectives such as cost and service quality. Existing optimization and reinforcement learning methods are effective for fixed formulations, but often depend on task-specific representations and constraint handling. We propose \emph{CityPlanner}, a sandbox-agent framework for executable urban planning. CityPlanner introduces \emph{UrbanSandbox}, a unified file-based environment where agents inspect task files, generate plans, run evaluators, and revise decisions based on executable feedback. To make learning tractable, we further propose atomic-task reinforcement learning, which decomposes long sandbox trajectories into \emph{BuildPlan} for initial construction and \emph{ImprovePlan} for feedback-based refinement. Experiments on a real-world benchmark show that CityPlanner consistently outperforms heuristic, task-specific RL, and general LLM-agent baselines. Ablations verify the contributions of UrbanSandbox, atomic-task RL, and iterative deployment. We release the code and dataset at https://anonymous.4open.science/r/co-agent-C1C8
Problem

Research questions and friction points this paper is trying to address.

Urban Planning
Spatial Optimization
Reinforcement Learning
Task-specific Representations
Constraint Handling
Innovation

Methods, ideas, or system contributions that make the work stand out.

CityPlanner
UrbanSandbox
atomic-task reinforcement learning
BuildPlan
ImprovePlan
🔎 Similar Papers
No similar papers found.
Wentao Zhang
Wentao Zhang
Institute of Physics, Chinese Academy of Sciences
photoemissionsuperconductivitycupratehtsctime-resolved
Jingyuan Wang
Jingyuan Wang
Beihang University
Data MiningSpatio-temporal Data MiningUrban ComputingCOVID-19
Z
Zetong Zhou
School of Computer Science and Engineering, Beihang University, Beijing, China; MIIT Key Laboratory of Data and Decision Intelligence, Beihang University, Beijing, China
Y
Yifan Yang
School of Computer Science and Engineering, Beihang University, Beijing, China; MIIT Key Laboratory of Data and Decision Intelligence, Beihang University, Beijing, China
W
Wenrui Wang
School of Computer Science and Engineering, Beihang University, Beijing, China; MIIT Key Laboratory of Data and Decision Intelligence, Beihang University, Beijing, China