ConsiSpace: 비디오 공간 추론을 위한 기하학적 일관성 학습의 중요성
ConsiSpace: Learning Geometric Consistency Matters for Video Spatial Reasoning
July 20, 2026
저자: Ting Huang, Zhenyu Zhang, Wenyuan Huang, Jian Yang, Hao Tang
cs.AI
초록
비디오 공간 추론은 내비게이션 지향 인지 및 장시간 비디오 질문 응답에 필수적이며, 모델은 변화하는 시점에서 장기간에 걸친 공간 관계를 추론해야 한다. 그러나 기존의 다중 모달 대규모 언어 모델(MLLM)은 대체로 의미 중심적이며, 중복된 비디오 관측으로부터 일관된 공간 증거를 신뢰성 있게 집계하지 못하여 비효율적이거나 불안정한 추론을 초래한다. 이러한 문제를 해결하기 위해, 우리는 기하학적 일관성 인식 프레임워크인 ConsiSpace를 제안한다. 이는 공간 일관성을 증거 조직 원칙이자 명시적인 사후 SFT 학습 신호로 활용한다. 우리는 암시적 증거 토큰과 명시적 기하학적 단서를 포함하는 기하학적 일관성 메모리(GCM)를 구축하고, 효율적인 조직 전략을 활용하여 작업 관련 공간 증거를 간결하게 보존한다. 또한, 지도 미세 조정 이후 통합 일관성 자기 지도 강화 학습(UC-SSRL)을 활용하여 응답, 측정 및 위상 일관성 보상을 통해 교차 뷰 안정성을 향상시킨다. 세 가지 공간 추론 벤치마크인 VSI-Bench, OSI-Bench 및 MMSI-Video-Bench에 대한 광범위한 실험에서 일관된 성능 향상을 보였으며, 가장 강력한 기준선 대비 평균 점수가 12.6포인트 향상되었다.
English
Video spatial reasoning is essential for navigation-oriented perception and long-video question answering, where models must infer spatial relations across long horizons under changing viewpoints. However, existing multimodal large language models (MLLMs) remain largely semantic-centric, and often fail to reliably aggregate consistent spatial evidence from redundant video observations, leading to inefficient or unstable reasoning. To address these issues, we propose ConsiSpace, a geometry-consistency-aware framework for geometry-sensitive video spatial reasoning that turns spatial consistency into both an evidence organization principle and an explicit post-SFT learning signal. We build a geometry-consistent memory (GCM) including implicit evidence tokens and explicit geometric cues, and leverage efficient organization strategies to compactly preserve task-related spatial evidence. Furthermore, we utilize unified consistency self-supervised reinforcement learning (UC-SSRL) after supervised fine-tuning to improve cross-view stability, with answer-, metric-, and topology-consistency rewards. Extensive experiments on three spatial-reasoning benchmarks, VSI-Bench, OSI-Bench, and MMSI-Video-Bench, show consistent gains, improving the average score by 12.6 points over the strongest baselines.