ChronoVision:透過潛在狀態重建的時間推理
ChronoVision: Temporal Reasoning via Latent State Reconstruction
August 6, 2026
作者: Yifan Shen, Jian Xu, Boyi Li, Yuner Zhang, Tianjiao Yu, Bingxuan Li, Houze Yang, Rushi Wang, Xu Cao
cs.AI
摘要
多模態大型語言模型擅長被動感知,但在需要多步驟時間推理的複雜視覺認知任務上卻表現吃力。這種退化主要源於基於語言的推理固有的模糊性,其往往無法準確表述連續的視覺變換。為了解決此問題,我們提出 ChronoVision,這是一個旨在將視覺邏輯與潛在圖像表徵對齊的多模態框架。在監督式微調期間,重建式視覺頭會預測最終轉換狀態的潛在表徵,同時 ROI 注意力定位模組透過語義跨度查詢,將模型聚焦於關鍵視覺證據。在後訓練階段,我們應用帶有隱式過程接地機制的強化學習,並由一個複合獎勵函數引導,該函數評估結果正確性、潛在過程對齊以及無監督視覺聚焦。此外,我們提出了 Vbvr-VQA,這是一個新穎的資料集,透過將影片推理重新表述為嚴格的影像排序任務,來評估時間追蹤能力。實驗表明,ChronoVision 在 Vbvr-VQA 上達到了最先進的效能,域內準確率為 74.8%,域外準確率為 71.6%,同時在極具挑戰性的跨域基準 IntPhys2 上也取得了 55.0% 的強勁準確率。
English
Multimodal large language models excel at passive perception but struggle with complex visual cognitive tasks requiring multi-step temporal reasoning. This degradation largely stems from the inherent ambiguity of language-based reasoning, which often fails to accurately articulate continuous visual transformations. To address this, we propose ChronoVision, a multimodal framework designed to align visual logic with latent imagery. During supervised fine-tuning, a Reconstructive Visual Head predicts the latent representation of the final transformed state, while an ROI Attention Locating module focuses the model on key visual evidence via semantic span queries. In post-training, we apply reinforcement learning with an implicit process grounding mechanism, guided by a composite reward function that evaluates outcome correctness, latent process alignment, and unsupervised visual focus. Furthermore, we introduce Vbvr-VQA, a novel dataset that evaluates temporal tracking by reformulating video reasoning into a strict image-ordering task. Experiments demonstrate that ChronoVision achieves state-of-the-art performance on Vbvr-VQA with 74.8% in-domain and 71.6% out-of-domain accuracy, alongside a strong 55.0% accuracy on IntPhys2, a highly challenging cross-domain benchmark.