ChatPaper.aiChatPaper

ChronoVision: 잠재 상태 재구성을 통한 시간적 추론

ChronoVision: Temporal Reasoning via Latent State Reconstruction

August 6, 2026
저자: Yifan Shen, Jian Xu, Boyi Li, Yuner Zhang, Tianjiao Yu, Bingxuan Li, Houze Yang, Rushi Wang, Xu Cao
cs.AI

초록

멀티모달 대규모 언어 모델은 수동적 지각에는 뛰어나지만, 다단계 시간적 추론을 요구하는 복잡한 시각적 인지 과제에는 어려움을 겪는다. 이러한 성능 저하의 주요 원인은 언어 기반 추론의 본질적 모호성에 있으며, 이는 연속적인 시각적 변환을 정확하게 표현하지 못하는 경우가 많다. 이 문제를 해결하기 위해 우리는 잠재 이미지와 시각적 논리를 정렬하도록 설계된 멀티모달 프레임워크인 ChronoVision을 제안한다. 지도 미세조정 단계에서 재구성적 시각 헤드(Reconstructive Visual Head)는 최종 변환 상태의 잠재 표현을 예측하고, ROI 주의 국소화 모듈(ROI Attention Locating module)은 의미적 범위 질의를 통해 핵심 시각적 증거에 모델이 집중하도록 유도한다. 사후 훈련 단계에서는 결과 정확성, 잠재 과정 정렬, 비지도 시각적 집중을 평가하는 복합 보상 함수에 기반한 암시적 과정 근거 부여 메커니즘을 갖춘 강화 학습을 적용한다. 또한, 비디오 추론을 엄격한 이미지 순서화 과제로 재구성하여 시간적 추적을 평가하는 새로운 데이터셋인 Vbvr-VQA를 도입한다. 실험 결과 ChronoVision은 Vbvr-VQA에서 도메인 내 정확도 74.8%, 도메인 외 정확도 71.6%로 최첨단 성능을 달성했으며, 매우 도전적인 교차 도메인 벤치마크인 IntPhys2에서도 55.0%의 높은 정확도를 달성함을 보여준다.
English
Multimodal large language models excel at passive perception but struggle with complex visual cognitive tasks requiring multi-step temporal reasoning. This degradation largely stems from the inherent ambiguity of language-based reasoning, which often fails to accurately articulate continuous visual transformations. To address this, we propose ChronoVision, a multimodal framework designed to align visual logic with latent imagery. During supervised fine-tuning, a Reconstructive Visual Head predicts the latent representation of the final transformed state, while an ROI Attention Locating module focuses the model on key visual evidence via semantic span queries. In post-training, we apply reinforcement learning with an implicit process grounding mechanism, guided by a composite reward function that evaluates outcome correctness, latent process alignment, and unsupervised visual focus. Furthermore, we introduce Vbvr-VQA, a novel dataset that evaluates temporal tracking by reformulating video reasoning into a strict image-ordering task. Experiments demonstrate that ChronoVision achieves state-of-the-art performance on Vbvr-VQA with 74.8% in-domain and 71.6% out-of-domain accuracy, alongside a strong 55.0% accuracy on IntPhys2, a highly challenging cross-domain benchmark.