超越視覺思維鏈:內化視覺思考以實現主動式影片推理
Beyond Visual CoT: Internalized Visual Thinking for Proactive Video Reasoning
August 16, 2026
作者: Xiaoyu Zhu, Xinke Deng, Suresh Taddewadikar, Arnab Kumar Mondal, Zhongyu Jiang, Ian Fasel, Joerg Liebelt
cs.AI
摘要
多模態大型語言模型日益使用視覺思維鏈(Visual CoT)來推理空間、時間與具身環境。透過生成中間推理圖像,Visual CoT 提供了直觀的視覺預見機制,但卻引入了大量的推論開銷,這在主動式影片推理中尤其造成問題。我們探討模型是否能在訓練期間學習視覺思考,同時在推論時直接進行推理。我們提出內化視覺思維(Internalized Visual Thinking, IVT),這是一個後訓練框架,在未標記影片上聯合優化文本預測與下一嵌入預測。給定部分觀測的影片,IVT 與目標文本答案一同預測未來幀的潛在表徵,促使模型捕捉運動、物體轉換、互動以及潛在意圖。在推論時,IVT 直接生成答案,無需合成或重新編碼未來幀。我們針對目標表徵、解碼器設計、預測時域、資料混合、訓練課程及預測目標進行了受控研究。IVT 在所有六種評估設定上均優於直接答案微調,同時保留了相同的推論路徑。與顯式 Visual CoT 相比,IVT 達到相當或更佳的效能,並將平均端到端延遲降低超過 5 倍。綜上所述,我們的研究結果表明,視覺思維鏈中所使用的推論時顯式像素空間生成,對於有效的主動式影片推理可能並非必要。預測性世界建模可以在訓練期間被內化,以產出既更準確又更高效的多模態推理器。
English
Multimodal large language models increasingly use visual chain-of-thought (Visual CoT) to reason about spatial, temporal, and embodied environments. By generating intermediate reasoning images, Visual CoT provides an intuitive mechanism for visual foresight but introduces substantial inference overhead, which is particularly problematic for proactive video reasoning. We ask whether models can learn to think visually during training while reasoning directly at inference. We introduce Internalized Visual Thinking (IVT), a post-training framework that jointly optimizes textual prediction and next-embedding prediction over unlabeled videos. Given a partially observed video, IVT predicts latent representations of future frames together with the target textual answer, encouraging the model to capture motion, object transitions, interactions, and latent intent. At inference, IVT generates the answer directly without synthesizing or re-encoding future frames. We conduct controlled studies across target representations, decoder designs, prediction horizons, data mixtures, training curricula, and predictive objectives. IVT improves over direct-answer fine-tuning on all six evaluation settings while retaining the same inference pathway. Compared with explicit Visual CoT, IVT achieves comparable or better performance and reduces average end-to-end latency by more than 5x. Together, our findings suggest that explicit pixel-space generation at inference time, as used in visual chain-of-thought, may not be necessary for effective proactive video reasoning. Predictive world modeling can be internalized during training to produce multimodal reasoners that are both more accurate and substantially more efficient.