ChatPaper.aiChatPaper

インターリーブされたマルチモーダル推論を統一的決定プロセスとして橋渡しする

Bridging Interleaved Multi-Modal Reasoning as a Unified Decision Process

July 4, 2026
著者: Zican Hu, Xuyang Hu, Yiming Liu, Zuwei Long, Wei Liu, Yunzhuo Hao, Jiawei Gu, Linjie Li, Yu Cheng, Zhenhong Sun, Weibo Gu, Xing Sun, Zhi Wang
cs.AI

要旨

統一マルチモーダルモデル(UMM)は、テキストと画像が混在した推論能力において有望な成果を示しているが、強化学習(RL)を用いてこのようなマルチターン生成を効果的に最適化することは依然として未解決の課題である。既存の手法ではRLをテキストステップにのみ適用し、画像生成は教師ありの代理タスクに委ねているため、ポリシー勾配が異種モダリティ間の全インタリーブ軌道に伝播しない。これにより、UMMにおけるRLの可能性はほとんど活用されていない。本論文では、マルチターンのテキスト-画像-テキスト推論を統一マルコフ決定過程(MDP)として捉えるシンプルなフレームワークBRAID(Bridging inteRleAved multI-modal reasoning as a unified Decision process)を導入する。これにより、単一の原理に基づくRL目的関数を通じて、テキスト生成と画像生成の統合的最適化が可能となる。BRAIDは共有の軌道レベルのアドバンテージを計算し、それをテキストトークンと画像デノイジングパスの両方に一貫して伝播させる。それぞれのモダリティ固有のポリシー勾配メカニズムによって最適化される。さらに、長期的なクレジット割り当てに対処するため、BRAIDは視覚言語モデル(VLM)判定器を採用し、各中間画像の推論有用性をスコアリングすることで、重要な視覚分岐において学習を強化する密なターンレベルのフィードバックを提供する。空間推論と視覚知覚のベンチマーク実験により、BRAIDは様々なベースラインを一貫して上回り、視覚的思考のガイダンスを備えた統一MDPの定式化が効果的なマルチモーダル推論に不可欠であることが確認された。
English
Unified multi-modal models (UMMs) have shown promising interleaved text-image reasoning capabilities, yet effectively optimizing such multi-turn generation via reinforcement learning (RL) remains an open challenge. Existing approaches apply RL exclusively to text steps, relegating image generation to supervised surrogates, preventing policy gradients from propagating through the full interleaved trajectory across heterogeneous modalities. This leaves the potential of RL for UMMs largely untapped. In the paper, we introduce BRAID (Bridging inteRleAved multI-modal reasoning as a unified Decision process), a simple framework that casts multi-turn text-image-text reasoning as a unified Markov decision process (MDP), enabling joint optimization of textual and visual generation via a single, principled RL objective. BRAID computes a shared trajectory-level advantage and propagates it coherently into both text tokens and image denoising paths, each optimized through its modality-native policy gradient mechanism. To further address long-horizon credit assignment, BRAID employs a vision-language model (VLM) judge that scores each intermediate image on its reasoning utility, supplying dense turn-level feedback to sharpen learning at critical visual branches. Experiments on spatial reasoning and visual perception benchmarks show that BRAID consistently outperforms various baselines, confirming that a unified MDP formulation with vision-thinking guidance is essential for effective multi-modal reasoning.