ChatPaper.aiChatPaper

視覚的CoTを超えて:能動的ビデオ推論のための内面化された視覚的思考

Beyond Visual CoT: Internalized Visual Thinking for Proactive Video Reasoning

August 16, 2026
著者: Xiaoyu Zhu, Xinke Deng, Suresh Taddewadikar, Arnab Kumar Mondal, Zhongyu Jiang, Ian Fasel, Joerg Liebelt
cs.AI

要旨

マルチモーダル大規模言語モデルは、空間的・時間的・身体的環境について推論するために、ますます視覚的思考連鎖(Visual CoT)を活用している。中間的な推論画像を生成することにより、Visual CoTは視覚的先見のための直感的なメカニズムを提供する一方で、 substantialな推論オーバーヘッドを導入し、これはプロアクティブな映像推論において特に問題となる。本研究では、モデルがトレーニング中に視覚的思考を学習し、推論時には直接的に推論できるかどうかを問う。我々は内在化視覚思考(Internalized Visual Thinking, IVT)を導入する。これは、ラベルなし映像に対してテキスト予測と次エンベディング予測を同時に最適化するポストトレーニングフレームワークである。部分的に観測された映像が与えられると、IVTは将来フレームの潜在表現をターゲットとなるテキスト回答とともに予測し、モデルが動き、物体の遷移、相互作用、および潜在的な意図を捕捉することを促進する。推論時には、IVTは将来フレームを合成または再エンコードすることなく、回答を直接生成する。我々は、ターゲット表現、デコーダ設計、予測ホライズン、データ混合、トレーニングカリキュラム、および予測目的関数にわたる対照実験を実施する。IVTは、6つの評価設定すべてにおいて直接回答ファインチューニングを上回り、同じ推論経路を維持する。明示的なVisual CoTと比較して、IVTは同等以上のパフォーマンスを達成し、平均エンドツーエンドレイテンシを5倍以上削減する。これらの知見は、視覚的思考連鎖で用いられる推論時の明示的なピクセル空間生成が、効果的なプロアクティブな映像推論には必ずしも必要ではない可能性を示唆する。予測的ワールドモデリングは、トレーニング中に内在化することで、より正確かつ大幅に効率的なマルチモーダル推論器を生成することができる。
English
Multimodal large language models increasingly use visual chain-of-thought (Visual CoT) to reason about spatial, temporal, and embodied environments. By generating intermediate reasoning images, Visual CoT provides an intuitive mechanism for visual foresight but introduces substantial inference overhead, which is particularly problematic for proactive video reasoning. We ask whether models can learn to think visually during training while reasoning directly at inference. We introduce Internalized Visual Thinking (IVT), a post-training framework that jointly optimizes textual prediction and next-embedding prediction over unlabeled videos. Given a partially observed video, IVT predicts latent representations of future frames together with the target textual answer, encouraging the model to capture motion, object transitions, interactions, and latent intent. At inference, IVT generates the answer directly without synthesizing or re-encoding future frames. We conduct controlled studies across target representations, decoder designs, prediction horizons, data mixtures, training curricula, and predictive objectives. IVT improves over direct-answer fine-tuning on all six evaluation settings while retaining the same inference pathway. Compared with explicit Visual CoT, IVT achieves comparable or better performance and reduces average end-to-end latency by more than 5x. Together, our findings suggest that explicit pixel-space generation at inference time, as used in visual chain-of-thought, may not be necessary for effective proactive video reasoning. Predictive world modeling can be internalized during training to produce multimodal reasoners that are both more accurate and substantially more efficient.