360CityArena:具身智能體的逼真虛擬城市導航基準
360CityArena: A Realistic Virtual Urban Navigation Benchmark for Embodied Agents
August 9, 2026
作者: Kenta Watanabe, Atsuyuki Miyai, Mizuki Takenawa, Kiyoharu Aizawa, Toshihiko Yamasaki
cs.AI
摘要
我們提出360CityArena,這是一個用於評估具身智能體在由360度影片構建之照片級寫實環境中進行城市探索能力的基準。現有的戶外基準要不是缺乏足夠的照片級寫實度,就是缺乏複雜度,導致與真實城市環境存在相當大的差距。360CityArena 建構於日本東京秋葉原地區的真實重建,使用了涵蓋85條街道的602段360度影片,並包含175個由人工精心設計的任務。它涵蓋三個任務類別:環境理解、路徑推理與空間推理,涵蓋城市探索所需的基本能力,例如定位、地標搜尋、路徑規劃與關係空間推理,從而能在真實城市場景中進行全面評估。我們使用基於最新大型多模態模型(LMM)的智能體進行評估,結果顯示即使是最強的模型 Gemini 2.5 Flash,其表現仍遠低於人類水準(人類:77.3% 對比 Gemini 2.5 Flash:17.1%),揭示了城市規模的具身導航與推理仍存在重大挑戰。360CityArena 為照片級寫實城市區域導航與空間推理提供了一個必要且具挑戰性的測試平台。
English
We present 360CityArena, a benchmark for evaluating the urban exploration capabilities of embodied agents within a photorealistic environment constructed from 360-degree videos. Existing outdoor benchmarks either lack sufficient photorealism or complexity, resulting in a considerable gap from real-world urban environments. 360CityArena is built on a realistic reconstruction of the Akihabara district in Tokyo, Japan, using 602 360-degree video segments covering 85 streets, and consists of 175 meticulously human-crafted tasks. It encompasses three task categories: Environment Understanding, Path Reasoning, and Spatial Reasoning, covering fundamental abilities required for urban exploration, such as localization, landmark search, path planning, and relational spatial reasoning, thereby enabling comprehensive evaluation in realistic urban scenes. Our evaluation using state-of-the-art LMM-based agents shows that even the strongest model, Gemini 2.5 Flash, performs far below human level (human: 77.3% vs. Gemini 2.5 Flash: 17.1%), revealing substantial challenges that remain in city-scale embodied navigation and reasoning. 360CityArena provides a necessary and challenging testbed for photorealistic urban-district navigation and spatial reasoning.