ChatPaper.aiChatPaper

GST-Bench: VLM이 비디오로부터 전역 공간 인식을 개발할 수 있는가?

GST-Bench: Can VLMs Develop Global Spatial Awareness from Video?

August 6, 2026
저자: Qifeng Zhang, Kaixiang Huang, Heng Dong, Huang Fang, Junting Chen, Junjie Zhu, Yonghang Chen, Zhiyu Zhang, Wei Li
cs.AI

초록

공간 지능은 구현 에이전트의 기본 요소이지만, 기존 벤치마크는 단일 또는 소수의 시점에서의 국소적 공간 지각에 초점을 맞추어, 연속적이고 장기간의 시각적 스트림에 대한 전역적 공간 인식을 간과하고 있다. 이러한 한계를 해결하기 위해, 우리는 비디오 이해에서의 전역적 공간 지능을 위한 VQA 벤치마크인 GST-Bench(Global-Spatial-Temporal Benchmark)를 소개한다. GST-Bench는 합성 생성된 6,790분의 비디오에서 도출된 인간 검증 질문으로 구성된다. 이 벤치마크는 모델이 입력 비디오에서 보이지 않는 새로운 시점에서의 정확한 공간 추론을 수행하고, 자기중심적 관찰을 전역적 탑다운 이미지에 매핑할 것을 요구한다. 22개의 최첨단 VLM에 대한 포괄적 평가는 모델과 인간 사이의 현저한 격차를 드러낸다: 가장 강력한 제로샷 모델은 42.68점에 그친 반면, 인간의 점수는 79.08점에 달한다. 이러한 격차의 원인을 조사하기 위해, 우리는 GST-Bench-Local을 구축하였고, 동일한 과제 구성 하에서 강력한 국소적 공간 이해 능력을 보임에도 불구하고 모델들이 장기간의 관찰을 전역적으로 일관된 장면 표현으로 통합하지 못한다는 것을 발견하였다. 또한, 이 과제에 대한 향후 연구를 촉진하기 위한 보완 자원으로서 전역적 공간 추론 데이터셋인 GST-Train을 제공한다.
English
Spatial intelligence is fundamental to embodied agents, yet existing benchmarks focus on local spatial perception from single or few viewpoints, overlooking global spatial awareness over continuous, long-horizon visual streams. To address this limitation, we introduce the Global-Spatial-Temporal Benchmark (GST-Bench), a VQA benchmark for global spatial intelligence in video understanding, comprising human-verified questions derived from 6,790 minutes of synthetically generated video. It requires models to perform accurate spatial inference from novel viewpoints unseen in the input video and to map egocentric observations onto global top-down images. A comprehensive evaluation of 22 state-of-the-art VLMs exposes a striking gap between models and humans: the strongest zero-shot model attains only 42.68, far below the human score of 79.08. To probe the cause of this gap, we construct GST-Bench-Local and find that models, despite strong local spatial understanding under the same task formulation, still fail to consolidate long-horizon observations into a globally consistent scene representation. We further provide GST-Train, a dataset for global spatial reasoning, as a complementary resource to facilitate future research on this challenge.