GST-Bench:视觉语言模型能否从视频中发展出全局空间感知?
GST-Bench: Can VLMs Develop Global Spatial Awareness from Video?
August 6, 2026
作者: Qifeng Zhang, Kaixiang Huang, Heng Dong, Huang Fang, Junting Chen, Junjie Zhu, Yonghang Chen, Zhiyu Zhang, Wei Li
cs.AI
摘要
空间智能是具身智能体的基础,然而现有基准仅关注从单一或少数视角进行的局部空间感知,忽视了在连续、长时间跨度的视觉流中的全局空间意识。为解决这一局限,我们提出了全局-空间-时间基准(GST-Bench),这是一个用于视频理解中全局空间智能的视觉问答基准,包含从6,790分钟合成视频中提取并经人工验证的问题。该基准要求模型从输入视频中未出现过的全新视角进行准确的空间推断,并将第一人称观察映射到全局俯视图。对22个最先进的视觉语言模型(VLM)的综合评估揭示了模型与人类之间的显著差距:最强的零样本模型仅取得42.68分,远低于人类的79.08分。为探究这一差距的根源,我们构建了GST-Bench-Local,发现尽管模型在同一任务设定下表现出较强的局部空间理解能力,但仍无法将长时间跨度的观察整合为全局一致的场景表征。我们进一步提供了GST-Train,一个面向全局空间推理的数据集,作为补充资源,以促进未来对这一挑战的研究。
English
Spatial intelligence is fundamental to embodied agents, yet existing benchmarks focus on local spatial perception from single or few viewpoints, overlooking global spatial awareness over continuous, long-horizon visual streams. To address this limitation, we introduce the Global-Spatial-Temporal Benchmark (GST-Bench), a VQA benchmark for global spatial intelligence in video understanding, comprising human-verified questions derived from 6,790 minutes of synthetically generated video. It requires models to perform accurate spatial inference from novel viewpoints unseen in the input video and to map egocentric observations onto global top-down images. A comprehensive evaluation of 22 state-of-the-art VLMs exposes a striking gap between models and humans: the strongest zero-shot model attains only 42.68, far below the human score of 79.08. To probe the cause of this gap, we construct GST-Bench-Local and find that models, despite strong local spatial understanding under the same task formulation, still fail to consolidate long-horizon observations into a globally consistent scene representation. We further provide GST-Train, a dataset for global spatial reasoning, as a complementary resource to facilitate future research on this challenge.