ChronoVision:潜在状態再構成による時間的推論
ChronoVision: Temporal Reasoning via Latent State Reconstruction
August 6, 2026
著者: Yifan Shen, Jian Xu, Boyi Li, Yuner Zhang, Tianjiao Yu, Bingxuan Li, Houze Yang, Rushi Wang, Xu Cao
cs.AI
要旨
マルチモーダル大規模言語モデルは受動的知覚には優れているものの、多段階の時間的推論を必要とする複雑な視覚認知タスクには苦戦している。この性能低下は、言語ベースの推論に本来備わる曖昧さに大きく起因しており、連続的な視覚変化を正確に言語化することがしばしば困難である。この問題に対処するため、我々は視覚的論理と潜在イメージを整合させるマルチモーダルフレームワークであるChronoVisionを提案する。教師ありファインチューニングでは、再構成型視覚ヘッドが最終変換状態の潜在表現を予測し、関心領域注意定位モジュールが意味的スパンクエリを介してモデルを重要な視覚的証拠に集中させる。ポストトレーニングでは、結果の正確性、潜在プロセスの整合性、教師なし視覚的焦点を評価する複合報酬関数に導かれた、暗黙的プロセス接地メカニズムを備えた強化学習を適用する。さらに、ビデオ推論を厳密な画像順序付けタスクに再構成することにより時間的追跡を評価する新しいデータセットVbvr-VQAを導入する。実験により、ChronoVisionはVbvr-VQAにおいてドメイン内精度74.8%、ドメイン外精度71.6%という最先端の性能を達成し、さらに高い難易度のクロスドメインベンチマークであるIntPhys2においても55.0%という堅牢な精度を達成することを実証する。
English
Multimodal large language models excel at passive perception but struggle with complex visual cognitive tasks requiring multi-step temporal reasoning. This degradation largely stems from the inherent ambiguity of language-based reasoning, which often fails to accurately articulate continuous visual transformations. To address this, we propose ChronoVision, a multimodal framework designed to align visual logic with latent imagery. During supervised fine-tuning, a Reconstructive Visual Head predicts the latent representation of the final transformed state, while an ROI Attention Locating module focuses the model on key visual evidence via semantic span queries. In post-training, we apply reinforcement learning with an implicit process grounding mechanism, guided by a composite reward function that evaluates outcome correctness, latent process alignment, and unsupervised visual focus. Furthermore, we introduce Vbvr-VQA, a novel dataset that evaluates temporal tracking by reformulating video reasoning into a strict image-ordering task. Experiments demonstrate that ChronoVision achieves state-of-the-art performance on Vbvr-VQA with 74.8% in-domain and 71.6% out-of-domain accuracy, alongside a strong 55.0% accuracy on IntPhys2, a highly challenging cross-domain benchmark.