UrbanGround:実スケール都市における局所的知覚から空間的エージェンシーへ

UrbanGround: From Local Perception to Spatial Agency in a Real-Scale City

August 27, 2026
著者: Tianjie Ju, Zheng Wu, Yueqing Sun, Yuhan Cui, Bobo Li, Shengqiong Wu, Pengzhou Cheng, Haodong Zhao, Zongru Wu, Xinbei Ma, Doris Zhang, Kunling Li, Mong-Li Lee, Wynne Hsu, Hao Fei, Qi Gu, Gongshen Liu, Zhuosheng Zhang
cs.AI

要旨

マルチモーダル大規模言語モデル(MLLM)は街並みの眺望を解釈できるが、都市における行動能力は、エージェントが移動を開始した後もそのような局所的証拠が有用であり続けるかどうかに依存する。本論文では、現在のMLLMエージェントが複雑な実規模の都市において、局所的な都市認識をどの程度信頼性の高い行動へと変換できるかを調査する。我々は、領土全体の3D地理空間データから構築した香港の物理的制約を備えたレプリカにおいて、この問いを検証可能にする最初のサンドボックスであるUrbanGroundを提案する。UrbanGroundは一人称視点からの閉ループ相互作用をサポートし、ナビゲーション用の対話型地図を提供する。エージェントは3D都市に直接入り、一人称視点で探索できる。我々の分析は、3つの研究課題を通じて空間的問題の拡大に従う。まず、能動的観察後に局所シーンを十分に接地させ、空間的質問に回答できるかどうかを検証する。次に、目的地がより遠く、より曖昧になるにつれて、その接地がナビゲーションを支援するかどうかを問う。最後に、結果として生じる行動が経路の利用可能性と歩行者の動きの変化に対して頑健であるかどうかを調べる。現代のMLLMエージェントは通常、視覚認識と短距離空間推論において有用な原子的能力を示す一方、方位認識と歩行者を考慮した移動は依然として信頼性が低い。その中心的欠陥は、拡張探索中に顕在化する。局所的能力が持続的な目標指向行動へと統合されず、誤差が効果的な修正なしに蓄積されるのである。我々はUrbanGroundが、複雑で開放的な都市環境において現在のMLLMエージェントがどの程度信頼性高く探索できるかについてのより広範な研究を支援することを期待する。
English
Multimodal large language models (MLLMs) can interpret a street view, but urban agency depends on whether such local evidence remains useful after the agent starts to move. In this paper, we investigate how far current MLLM agents can turn local urban perception into reliable action in a complicated real-scale city. We propose UrbanGround, the first sandbox to make this question testable in a physically constrained replica of Hong Kong built from territory-wide 3D geospatial data. UrbanGround supports closed-loop interaction from a first-person view and provides an interactive map for navigation. Agents can directly enter the 3D city and explore from a first-person view. Our analysis follows the growth of the spatial problem through three research questions. We first test whether an agent can ground a local scene well enough to answer spatial questions after active observation. Then we ask whether that grounding supports navigation as destinations become farther away and less explicit. Finally, we examine whether the resulting behavior survives changes in route availability and pedestrian motion. Contemporary MLLM agents usually show useful atomic abilities in visual recognition and short-range spatial reasoning, while orientation and pedestrian-aware movement remain unreliable. Their central failure emerges over extended exploration, where local abilities do not compose into sustained goal-directed behavior and errors accumulate without effective correction. We hope UrbanGround will support broader study of how far current MLLM agents can explore reliably in complex, open-ended urban environments.
PDF692August 29, 2026