ChatPaper.aiChatPaper

ConsiSpace: ビデオ空間推論における幾何的一貫性学習の重要性

ConsiSpace: Learning Geometric Consistency Matters for Video Spatial Reasoning

July 20, 2026
著者: Ting Huang, Zhenyu Zhang, Wenyuan Huang, Jian Yang, Hao Tang
cs.AI

要旨

ビデオ空間推論は、ナビゲーション指向の知覚や長尺ビデオ質問応答において不可欠であり、モデルは視点が変化する長期的な時間スパンにわたって空間関係を推論する必要がある。しかし、既存のマルチモーダル大規模言語モデル(MLLM)は依然としてセマンティクス中心であり、冗長なビデオ観測から一貫した空間的証拠を確実に集約できず、非効率または不安定な推論を引き起こすことが多い。これらの問題に対処するため、我々はConsiSpaceを提案する。これは、幾何学的一貫性を考慮したフレームワークであり、空間的一貫性を証拠整理の原理として、また教師付きファインチューニング(SFT)後の明示的な学習信号として活用する。我々は、暗黙的証拠トークンと明示的な幾何学的手がかりを含む幾何一貫メモリ(GCM)を構築し、効率的な整理戦略を活用してタスク関連の空間的証拠をコンパクトに保持する。さらに、教師付きファインチューニング後に統一的一貫性自己教師あり強化学習(UC-SSRL)を適用し、回答、メトリクス、トポロジーの一貫性に関する報酬を用いて、ビュー間の安定性を向上させる。3つの空間推論ベンチマーク(VSI-Bench、OSI-Bench、MMSI-Video-Bench)における広範な実験により、一貫した改善が示され、最強のベースラインと比較して平均スコアが12.6ポイント向上した。
English
Video spatial reasoning is essential for navigation-oriented perception and long-video question answering, where models must infer spatial relations across long horizons under changing viewpoints. However, existing multimodal large language models (MLLMs) remain largely semantic-centric, and often fail to reliably aggregate consistent spatial evidence from redundant video observations, leading to inefficient or unstable reasoning. To address these issues, we propose ConsiSpace, a geometry-consistency-aware framework for geometry-sensitive video spatial reasoning that turns spatial consistency into both an evidence organization principle and an explicit post-SFT learning signal. We build a geometry-consistent memory (GCM) including implicit evidence tokens and explicit geometric cues, and leverage efficient organization strategies to compactly preserve task-related spatial evidence. Furthermore, we utilize unified consistency self-supervised reinforcement learning (UC-SSRL) after supervised fine-tuning to improve cross-view stability, with answer-, metric-, and topology-consistency rewards. Extensive experiments on three spatial-reasoning benchmarks, VSI-Bench, OSI-Bench, and MMSI-Video-Bench, show consistent gains, improving the average score by 12.6 points over the strongest baselines.