UrbanGround:從在地感知到真實尺度城市中的空間能動性

UrbanGround: From Local Perception to Spatial Agency in a Real-Scale City

August 27, 2026
作者: Tianjie Ju, Zheng Wu, Yueqing Sun, Yuhan Cui, Bobo Li, Shengqiong Wu, Pengzhou Cheng, Haodong Zhao, Zongru Wu, Xinbei Ma, Doris Zhang, Kunling Li, Mong-Li Lee, Wynne Hsu, Hao Fei, Qi Gu, Gongshen Liu, Zhuosheng Zhang
cs.AI

摘要

多模態大型語言模型(MLLMs)能夠解讀街道景觀,但城市行動能力取決於當智能體開始移動後,這些局部證據是否仍然有用。在本文中,我們探討當前的 MLLM 智能體如何在複雜的真實規模城市中,將局部的城市感知轉化為可靠的行動。我們提出 UrbanGround,這是首個使此問題得以測試的沙盒,其環境是根據覆蓋全港的三維地理空間數據所建構、且受物理約束的香港複製城市。UrbanGround 支援第一人稱視角的閉環互動,並提供互動式地圖以供導航。智能體可以直接進入三維城市,並以第一人稱視角進行探索。我們的分析循著空間問題的擴展,透過三個研究問題逐步推進。首先,我們測試智能體在主動觀察後,是否能充分錨定局部場景以回答空間問題。接著,我們探討當目的地變得更遠、更不明確時,這種錨定能力是否足以支援導航。最後,我們檢視所產生的行為能否在路線可用性與行人運動的變化下維持。當代 MLLM 智能體通常在視覺辨識與短距離空間推理方面展現有用的基礎能力,但方向感與具行人感知的移動仍不可靠。其核心失敗出現在長時間的探索過程中:局部能力無法組合成持續的目標導向行為,且錯誤在缺乏有效修正下持續累積。我們希望 UrbanGround 能促進更廣泛的研究,探討當前 MLLM 智能體在複雜、開放的都市環境中能可靠探索多遠。
English
Multimodal large language models (MLLMs) can interpret a street view, but urban agency depends on whether such local evidence remains useful after the agent starts to move. In this paper, we investigate how far current MLLM agents can turn local urban perception into reliable action in a complicated real-scale city. We propose UrbanGround, the first sandbox to make this question testable in a physically constrained replica of Hong Kong built from territory-wide 3D geospatial data. UrbanGround supports closed-loop interaction from a first-person view and provides an interactive map for navigation. Agents can directly enter the 3D city and explore from a first-person view. Our analysis follows the growth of the spatial problem through three research questions. We first test whether an agent can ground a local scene well enough to answer spatial questions after active observation. Then we ask whether that grounding supports navigation as destinations become farther away and less explicit. Finally, we examine whether the resulting behavior survives changes in route availability and pedestrian motion. Contemporary MLLM agents usually show useful atomic abilities in visual recognition and short-range spatial reasoning, while orientation and pedestrian-aware movement remain unreliable. Their central failure emerges over extended exploration, where local abilities do not compose into sustained goal-directed behavior and errors accumulate without effective correction. We hope UrbanGround will support broader study of how far current MLLM agents can explore reliably in complex, open-ended urban environments.
PDF692August 29, 2026