GST-Bench:VLMsは動画からグローバル空間認識を獲得できるか?
GST-Bench: Can VLMs Develop Global Spatial Awareness from Video?
August 6, 2026
著者: Qifeng Zhang, Kaixiang Huang, Heng Dong, Huang Fang, Junting Chen, Junjie Zhu, Yonghang Chen, Zhiyu Zhang, Wei Li
cs.AI
要旨
空間知能は身体性エージェントにとって基本的な能力であるが、既存のベンチマークは単一または少数の視点からの局所的空間知覚に焦点を当てており、連続的かつ長期的な視覚ストリームにわたるグローバルな空間認識を見落としている。この限界に対処するため、我々はビデオ理解におけるグローバル空間知能のためのVQAベンチマークであるGlobal-Spatial-Temporal Benchmark(GST-Bench)を紹介する。GST-Benchは、6,790分の合成生成ビデオから導出された人間検証済みの質問で構成される。これは、モデルが入力ビデオには存在しない新規視点から正確な空間推論を行い、自己中心的観察をグローバルな俯瞰画像にマッピングすることを要求する。22の最先端の視覚言語モデル(VLM)に対する包括的な評価は、モデルと人間の間の顕著なギャップを浮き彫りにする:最強のゼロショットモデルは42.68しか達成できず、人間のスコア79.08を大きく下回る。このギャップの原因を探るため、我々はGST-Bench-Localを構築し、モデルが同一のタスク形式の下で強力な局所的空間理解を示すにもかかわらず、長期的な観察をグローバルに一貫したシーン表現へと統合することに依然として失敗することを見出した。さらに、この課題に関する将来の研究を促進するための補完的リソースとして、グローバル空間推論のためのデータセットGST-Trainを提供する。
English
Spatial intelligence is fundamental to embodied agents, yet existing benchmarks focus on local spatial perception from single or few viewpoints, overlooking global spatial awareness over continuous, long-horizon visual streams. To address this limitation, we introduce the Global-Spatial-Temporal Benchmark (GST-Bench), a VQA benchmark for global spatial intelligence in video understanding, comprising human-verified questions derived from 6,790 minutes of synthetically generated video. It requires models to perform accurate spatial inference from novel viewpoints unseen in the input video and to map egocentric observations onto global top-down images. A comprehensive evaluation of 22 state-of-the-art VLMs exposes a striking gap between models and humans: the strongest zero-shot model attains only 42.68, far below the human score of 79.08. To probe the cause of this gap, we construct GST-Bench-Local and find that models, despite strong local spatial understanding under the same task formulation, still fail to consolidate long-horizon observations into a globally consistent scene representation. We further provide GST-Train, a dataset for global spatial reasoning, as a complementary resource to facilitate future research on this challenge.