비주얼 CoT를 넘어서: 능동적 비디오 추론을 위한 내면화된 시각적 사고
Beyond Visual CoT: Internalized Visual Thinking for Proactive Video Reasoning
August 16, 2026
저자: Xiaoyu Zhu, Xinke Deng, Suresh Taddewadikar, Arnab Kumar Mondal, Zhongyu Jiang, Ian Fasel, Joerg Liebelt
cs.AI
초록
다중모달 대규모 언어 모델은 공간적, 시간적, 그리고 체화된 환경에 대해 추론하기 위해 시각적 사고 사슬(Visual CoT)을 점점 더 많이 활용하고 있다. 중간 추론 이미지를 생성함으로써, Visual CoT는 시각적 예측을 위한 직관적 메커니즘을 제공하지만 상당한 추론 오버헤드를 초래하며, 이는 능동적 비디오 추론에서 특히 문제가 된다. 본 연구는 모델이 추론 시에는 직접 추론하면서 훈련 중에는 시각적으로 사고하는 법을 학습할 수 있는지 묻는다. 본 연구는 내재화된 시각적 사고(Internalized Visual Thinking, IVT)를 제안하는데, 이는 레이블이 없는 비디오에 대해 텍스트 예측과 다음 임베딩 예측을 공동으로 최적화하는 사후 훈련 프레임워크이다. 부분적으로 관찰된 비디오가 주어지면, IVT는 대상 텍스트 답변과 함께 미래 프레임의 잠재 표현을 예측하여, 모델이 움직임, 객체 전환, 상호작용, 그리고 잠재적 의도를 포착하도록 유도한다. 추론 시에 IVT는 미래 프레임을 합성하거나 재인코딩하지 않고 답변을 직접 생성한다. 본 연구는 대상 표현, 디코더 설계, 예측 지평, 데이터 혼합, 훈련 커리큘럼, 그리고 예측 목적 함수에 걸쳐 통제된 연구를 수행한다. IVT는 모든 여섯 가지 평가 설정에서 직접 답변 미세 조정보다 우수한 성능을 보이면서도 동일한 추론 경로를 유지한다. 명시적 Visual CoT와 비교하여, IVT는 유사하거나 더 나은 성능을 달성하고 평균 종단 간 지연 시간을 5배 이상 줄인다. 종합하면, 본 연구의 결과는 시각적 사고 사슬에서 사용되는 추론 시점의 명시적 픽셀 공간 생성이 효과적인 능동적 비디오 추론에 반드시 필요한 것은 아님을 시사한다. 예측적 세계 모델링은 훈련 중에 내재화되어 더 정확하면서도 훨씬 더 효율적인 다중모달 추론기를 만들어낼 수 있다.
English
Multimodal large language models increasingly use visual chain-of-thought (Visual CoT) to reason about spatial, temporal, and embodied environments. By generating intermediate reasoning images, Visual CoT provides an intuitive mechanism for visual foresight but introduces substantial inference overhead, which is particularly problematic for proactive video reasoning. We ask whether models can learn to think visually during training while reasoning directly at inference. We introduce Internalized Visual Thinking (IVT), a post-training framework that jointly optimizes textual prediction and next-embedding prediction over unlabeled videos. Given a partially observed video, IVT predicts latent representations of future frames together with the target textual answer, encouraging the model to capture motion, object transitions, interactions, and latent intent. At inference, IVT generates the answer directly without synthesizing or re-encoding future frames. We conduct controlled studies across target representations, decoder designs, prediction horizons, data mixtures, training curricula, and predictive objectives. IVT improves over direct-answer fine-tuning on all six evaluation settings while retaining the same inference pathway. Compared with explicit Visual CoT, IVT achieves comparable or better performance and reduces average end-to-end latency by more than 5x. Together, our findings suggest that explicit pixel-space generation at inference time, as used in visual chain-of-thought, may not be necessary for effective proactive video reasoning. Predictive world modeling can be internalized during training to produce multimodal reasoners that are both more accurate and substantially more efficient.