UrbanGround: 실제 규모의 도시에서 지역적 인지로부터 공간적 행위성으로
UrbanGround: From Local Perception to Spatial Agency in a Real-Scale City
August 27, 2026
저자: Tianjie Ju, Zheng Wu, Yueqing Sun, Yuhan Cui, Bobo Li, Shengqiong Wu, Pengzhou Cheng, Haodong Zhao, Zongru Wu, Xinbei Ma, Doris Zhang, Kunling Li, Mong-Li Lee, Wynne Hsu, Hao Fei, Qi Gu, Gongshen Liu, Zhuosheng Zhang
cs.AI
초록
멀티모달 대규모 언어 모델(MLLM)은 거리 뷰를 해석할 수 있지만, 도시적 행위성(urban agency)은 에이전트가 이동을 시작한 후에도 그러한 지역적 증거가 유용하게 유지되는지에 달려 있다. 본 논문에서는 현재의 MLLM 에이전트가 복잡한 실규모 도시에서 지역적 도시 인식을 얼마나 신뢰할 수 있는 행동으로 전환할 수 있는지를 조사한다. 우리는 홍콩 전역의 3차원 지리공간 데이터로 구축된 물리적 제약이 있는 홍콩 복제 환경에서 이 질문을 시험 가능하게 만드는 최초의 샌드박스인 UrbanGround를 제안한다. UrbanGround는 1인칭 시점의 폐루프 상호작용을 지원하며 내비게이션을 위한 대화형 지도를 제공한다. 에이전트는 3D 도시에 직접 진입하여 1인칭 시점으로 탐색할 수 있다. 우리의 분석은 세 가지 연구 질문을 통해 공간 문제의 성장을 추적한다. 먼저, 에이전트가 능동적 관찰 후 공간 질문에 답할 수 있을 만큼 지역 장면을 충분히 근거 지을 수 있는지 시험한다. 다음으로, 목적지가 더 멀어지고 덜 명시적이 될 때 그러한 근거 설정이 내비게이션을 지원하는지 묻는다. 마지막으로, 그 결과적인 행동이 경로 가용성과 보행자 이동의 변화에도 유지되는지 검토한다. 현행 MLLM 에이전트는 일반적으로 시각적 인식과 단거리 공간 추론에서 유용한 원자적 능력을 보여주지만, 방향 감각과 보행자 인식 이동은 여전히 신뢰할 수 없다. 그들의 핵심 실패는 확장된 탐색 과정에서 드러나는데, 지역적 능력이 지속적인 목표 지향 행동으로 구성되지 못하고 오류가 효과적 교정 없이 누적된다. 우리는 UrbanGround가 복잡하고 개방적인 도시 환경에서 현재 MLLM 에이전트가 얼마나 멀리 신뢰할 수 있게 탐색할 수 있는지에 대한 더 폭넓은 연구를 지원하기를 기대한다.
English
Multimodal large language models (MLLMs) can interpret a street view, but urban agency depends on whether such local evidence remains useful after the agent starts to move. In this paper, we investigate how far current MLLM agents can turn local urban perception into reliable action in a complicated real-scale city. We propose UrbanGround, the first sandbox to make this question testable in a physically constrained replica of Hong Kong built from territory-wide 3D geospatial data. UrbanGround supports closed-loop interaction from a first-person view and provides an interactive map for navigation. Agents can directly enter the 3D city and explore from a first-person view. Our analysis follows the growth of the spatial problem through three research questions. We first test whether an agent can ground a local scene well enough to answer spatial questions after active observation. Then we ask whether that grounding supports navigation as destinations become farther away and less explicit. Finally, we examine whether the resulting behavior survives changes in route availability and pedestrian motion. Contemporary MLLM agents usually show useful atomic abilities in visual recognition and short-range spatial reasoning, while orientation and pedestrian-aware movement remain unreliable. Their central failure emerges over extended exploration, where local abilities do not compose into sustained goal-directed behavior and errors accumulate without effective correction. We hope UrbanGround will support broader study of how far current MLLM agents can explore reliably in complex, open-ended urban environments.