城市地面:从在地感知到真实尺度城市中的空间能动性

UrbanGround: From Local Perception to Spatial Agency in a Real-Scale City

August 27, 2026
作者: Tianjie Ju, Zheng Wu, Yueqing Sun, Yuhan Cui, Bobo Li, Shengqiong Wu, Pengzhou Cheng, Haodong Zhao, Zongru Wu, Xinbei Ma, Doris Zhang, Kunling Li, Mong-Li Lee, Wynne Hsu, Hao Fei, Qi Gu, Gongshen Liu, Zhuosheng Zhang
cs.AI

摘要

多模态大语言模型(MLLMs)能够解读街景,但城市行动能力取决于这些局部证据在智能体开始移动后是否仍然有效。本文研究当前MLLM智能体在复杂的真实规模城市中能将多少局部城市感知转化为可靠行动。我们提出UrbanGround,这是首个基于香港全域三维地理空间数据、在具有物理约束的香港复刻环境中使该问题可测试的沙盒平台。UrbanGround支持第一人称视角下的闭环交互,并提供交互式地图用于导航。智能体可以直接进入三维城市,以第一人称视角进行探索。我们的分析通过三个研究问题追踪空间问题的增长过程。首先,我们测试智能体在主动观察后,是否能够充分理解局部场景以回答空间问题。然后,我们考察当目的地越来越远且越来越不明确时,这种理解能否支撑导航。最后,我们检验由此产生的行为在路径可用性和行人运动发生变化时是否依旧成立。当前MLLM智能体通常在视觉识别和短距离空间推理上表现出有用的原子能力,但方向感知和行人感知运动仍不可靠。它们的核心失败出现在长时间探索中:局部能力无法组合成持续的目标导向行为,错误不断累积且得不到有效纠正。我们希望UrbanGround能够支持更广泛的研究,探讨当前MLLM智能体在复杂、开放的城市环境中究竟能可靠地探索多远。
English
Multimodal large language models (MLLMs) can interpret a street view, but urban agency depends on whether such local evidence remains useful after the agent starts to move. In this paper, we investigate how far current MLLM agents can turn local urban perception into reliable action in a complicated real-scale city. We propose UrbanGround, the first sandbox to make this question testable in a physically constrained replica of Hong Kong built from territory-wide 3D geospatial data. UrbanGround supports closed-loop interaction from a first-person view and provides an interactive map for navigation. Agents can directly enter the 3D city and explore from a first-person view. Our analysis follows the growth of the spatial problem through three research questions. We first test whether an agent can ground a local scene well enough to answer spatial questions after active observation. Then we ask whether that grounding supports navigation as destinations become farther away and less explicit. Finally, we examine whether the resulting behavior survives changes in route availability and pedestrian motion. Contemporary MLLM agents usually show useful atomic abilities in visual recognition and short-range spatial reasoning, while orientation and pedestrian-aware movement remain unreliable. Their central failure emerges over extended exploration, where local abilities do not compose into sustained goal-directed behavior and errors accumulate without effective correction. We hope UrbanGround will support broader study of how far current MLLM agents can explore reliably in complex, open-ended urban environments.
PDF692August 29, 2026