360CityArena:身体化エージェントのための現実的な仮想都市ナビゲーションベンチマーク
360CityArena: A Realistic Virtual Urban Navigation Benchmark for Embodied Agents
August 9, 2026
著者: Kenta Watanabe, Atsuyuki Miyai, Mizuki Takenawa, Kiyoharu Aizawa, Toshihiko Yamasaki
cs.AI
要旨
本稿では、360度ビデオから構築されたフォトリアリスティックな環境において、身体化エージェントの都市探索能力を評価するためのベンチマークである360CityArenaを提案する。既存の屋外ベンチマークは、十分なフォトリアリズムまたは複雑さを欠いており、現実世界の都市環境との間にかなりのギャップが生じている。360CityArenaは、日本の東京・秋葉原地区を現実的に再構築したものであり、85の通りをカバーする602本の360度ビデオセグメントを用いて構築され、175の細心の注意を払って人手で作成されたタスクから構成される。これには、環境理解、経路推論、空間推論という3つのタスクカテゴリが含まれ、位置特定、ランドマーク探索、経路計画、関係的空間推論など、都市探索に必要な基本的な能力をカバーすることで、現実的な都市シーンにおける包括的な評価を可能にする。最先端のLMMベースのエージェントを用いた評価では、最も強力なモデルであるGemini 2.5 Flashでも、人間のレベルをはるかに下回る性能しか示さず(人間:77.3%、Gemini 2.5 Flash:17.1%)、都市規模の身体化ナビゲーションと推論には依然として大きな課題が残されていることが明らかになった。360CityArenaは、フォトリアリスティックな市街地ナビゲーションと空間推論のための、必要かつ挑戦的なテストベッドを提供する。
English
We present 360CityArena, a benchmark for evaluating the urban exploration capabilities of embodied agents within a photorealistic environment constructed from 360-degree videos. Existing outdoor benchmarks either lack sufficient photorealism or complexity, resulting in a considerable gap from real-world urban environments. 360CityArena is built on a realistic reconstruction of the Akihabara district in Tokyo, Japan, using 602 360-degree video segments covering 85 streets, and consists of 175 meticulously human-crafted tasks. It encompasses three task categories: Environment Understanding, Path Reasoning, and Spatial Reasoning, covering fundamental abilities required for urban exploration, such as localization, landmark search, path planning, and relational spatial reasoning, thereby enabling comprehensive evaluation in realistic urban scenes. Our evaluation using state-of-the-art LMM-based agents shows that even the strongest model, Gemini 2.5 Flash, performs far below human level (human: 77.3% vs. Gemini 2.5 Flash: 17.1%), revealing substantial challenges that remain in city-scale embodied navigation and reasoning. 360CityArena provides a necessary and challenging testbed for photorealistic urban-district navigation and spatial reasoning.