將交錯多模態推理橋接為統一決策過程
Bridging Interleaved Multi-Modal Reasoning as a Unified Decision Process
July 4, 2026
作者: Zican Hu, Xuyang Hu, Yiming Liu, Zuwei Long, Wei Liu, Yunzhuo Hao, Jiawei Gu, Linjie Li, Yu Cheng, Zhenhong Sun, Weibo Gu, Xing Sun, Zhi Wang
cs.AI
摘要
統一多模態模型(UMMs)在交錯式文字-圖像推理方面展現了優異能力,然而如何透過強化學習(RL)有效優化此類多輪生成仍是一項開放性挑戰。現有方法僅將RL應用於文字步驟,而將圖像生成交由監督式替代模型處理,使得策略梯度無法跨越異質模態的完整交錯軌跡進行傳播,導致RL在UMMs中的潛力未被充分發掘。本文提出BRAID(Bridging inteRleAved multI-modal reasoning as a unified Decision process),這是一個簡潔框架,將多輪文字-圖像-文字推理視為統一的馬可夫決策過程(MDP),從而透過單一、原則性的RL目標聯合優化文字與視覺生成。BRAID計算共享的軌跡層級優勢,並將其一致地傳播至文字令牌與圖像去雜訊路徑,各模態透過其原生策略梯度機制進行優化。為進一步處理長時程信用分配問題,BRAID採用視覺語言模型(VLM)裁判,根據各中間圖像的推理效用進行評分,提供密集的輪次級回饋,以強化關鍵視覺分支的學習。在空間推理與視覺感知基準測試上的實驗顯示,BRAID consistently 優於多種基線方法,證實以視覺思維引導的統一MDP公式對有效多模態推理至關重要。
English
Unified multi-modal models (UMMs) have shown promising interleaved text-image reasoning capabilities, yet effectively optimizing such multi-turn generation via reinforcement learning (RL) remains an open challenge. Existing approaches apply RL exclusively to text steps, relegating image generation to supervised surrogates, preventing policy gradients from propagating through the full interleaved trajectory across heterogeneous modalities. This leaves the potential of RL for UMMs largely untapped. In the paper, we introduce BRAID (Bridging inteRleAved multI-modal reasoning as a unified Decision process), a simple framework that casts multi-turn text-image-text reasoning as a unified Markov decision process (MDP), enabling joint optimization of textual and visual generation via a single, principled RL objective. BRAID computes a shared trajectory-level advantage and propagates it coherently into both text tokens and image denoising paths, each optimized through its modality-native policy gradient mechanism. To further address long-horizon credit assignment, BRAID employs a vision-language model (VLM) judge that scores each intermediate image on its reasoning utility, supplying dense turn-level feedback to sharpen learning at critical visual branches. Experiments on spatial reasoning and visual perception benchmarks show that BRAID consistently outperforms various baselines, confirming that a unified MDP formulation with vision-thinking guidance is essential for effective multi-modal reasoning.