ChatPaper.aiChatPaper

ConsiSpace:學習幾何一致性對影片空間推理至關重要

ConsiSpace: Learning Geometric Consistency Matters for Video Spatial Reasoning

July 20, 2026
作者: Ting Huang, Zhenyu Zhang, Wenyuan Huang, Jian Yang, Hao Tang
cs.AI

摘要

视频空间推理对于面向导航的感知与长视频问答至关重要,在此类任务中,模型必须在视角不断变化的长时程范围内推断空间关系。然而,现有的大语言多模态模型仍以语义为中心,往往无法从冗余的视频观测中可靠地聚合一致的空间证据,导致推理效率低下或不稳定。为解决这些问题,我们提出ConsiSpace,一种基于几何一致性感知的框架,专门用于几何敏感的视频空间推理,将空间一致性同时转化为证据组织原则与显式的后SFT学习信号。我们构建了包含隐式证据标记与显式几何线索的几何一致记忆模块,并利用高效的组织策略紧凑地保留与任务相关的空间证据。此外,我们采用监督微调后的统一一致性自监督强化学习,通过答案一致性、度量一致性与拓扑一致性奖励,提升跨视角稳定性。在三个空间推理基准测试(VSI-Bench、OSI-Bench与MMSI-Video-Bench)上的广泛实验显示,该方法相较于最强基线模型平均提升了12.6分,取得了持续的性能改善。
English
Video spatial reasoning is essential for navigation-oriented perception and long-video question answering, where models must infer spatial relations across long horizons under changing viewpoints. However, existing multimodal large language models (MLLMs) remain largely semantic-centric, and often fail to reliably aggregate consistent spatial evidence from redundant video observations, leading to inefficient or unstable reasoning. To address these issues, we propose ConsiSpace, a geometry-consistency-aware framework for geometry-sensitive video spatial reasoning that turns spatial consistency into both an evidence organization principle and an explicit post-SFT learning signal. We build a geometry-consistent memory (GCM) including implicit evidence tokens and explicit geometric cues, and leverage efficient organization strategies to compactly preserve task-related spatial evidence. Furthermore, we utilize unified consistency self-supervised reinforcement learning (UC-SSRL) after supervised fine-tuning to improve cross-view stability, with answer-, metric-, and topology-consistency rewards. Extensive experiments on three spatial-reasoning benchmarks, VSI-Bench, OSI-Bench, and MMSI-Video-Bench, show consistent gains, improving the average score by 12.6 points over the strongest baselines.