ChatPaper.aiChatPaper

潜在を通じて推論せよ!潜在視覚推論を必然化する

Reason Through the Latent! Making Latent Visual Reasoning Necessary

September 6, 2026
著者: Suhyeong Park, Junha Jung, Jaewoo Kang
cs.AI

要旨

潜在的視覚推論は、明示的なテキストによる思考連鎖ではなく、隠れ状態計算を通じてマルチモーダル推論を行うことを目指している。しかし、視覚情報が潜在状態に存在することは、モデルが答えを生成する際に実際にその状態に依存していることを意味しない。特に、代替の画像条件付き経路が利用可能な場合にはそうである。我々は因果的視覚再帰推論(CVRR)を導入する。これは事前学習済みの視覚能力を保持しつつ、再帰的計算を予測への必須の画像条件付き経路とする。CVRRは、事前学習済み視覚言語モデルが画像を取り込んだ後、質問の隠れ状態から再帰を初期化し、その後、同じ固定された視覚的証拠を再読しながらこの状態を繰り返し更新する。デコード前に、視覚状態と元のマルチモーダルKVキャッシュが除去され、最終的な再帰状態のみが画像条件付き情報を答えに伝える。V^*、MMVP、BLINK、MME-RealWorld-Liteの各ベンチマークにおいて、CVRRはこの厳密なインターフェースの下で強い性能を維持する一方、互換性のある潜在推論器は同じ制約の下で再訓練しても同等の視覚能力を回復できない。因果的介入はさらに、質問を固定したときに予測が再帰的内容に敏感であり続けること、および持続的な視覚的証拠が再帰的軌道を因果的に修正することを示す。これらの結果は、潜在的な情報量と、実際に予測に用いられる潜在計算とを区別する。
English
Latent visual reasoning aims to perform multimodal reasoning through hidden-state computation rather than explicit textual chains of thought. However, visual information being present in a latent state does not imply that the model actually relies on that state when producing its answer, especially when alternative image-conditioned paths remain available. We introduce Causal Visual Recurrent Reasoning (CVRR), which preserves pretrained visual competence while making recurrent computation the required image-conditioned path to prediction. CVRR initializes recurrence from the question hidden state after the pretrained vision-language model has incorporated the image, then repeatedly updates this state while re-reading the same fixed visual evidence. Before decoding, visual states and the original multimodal KV cache are removed so that only the final recurrent state carries image-conditioned information to the answer. Across the V^*, MMVP, BLINK, and MME-RealWorld-Lite benchmarks, CVRR retains strong performance under this strict interface, while compatible latent reasoners fail to recover comparable visual competence even when retrained under the same constraint. Causal interventions further show that predictions remain sensitive to recurrent content when the question is held fixed, and that persistent visual evidence causally revises the recurrent trajectory. These results distinguish latent informativeness from latent computation that is actually used for prediction.