DeepVoyager-VL:激勵長程多模態智能體進行視覺在環搜尋
DeepVoyager-VL: Incentivizing Vision-in-the-Loop Search for Long-Horizon Multimodal Agents
August 3, 2026
作者: Huanyao Zhang, Jiepeng Zhou, Runhao Zhao, Yanzhe Shan, Jiaoyang Chen, Bowen Zhou, Bo Li, Fang Wang, Jialong Wu, Zhengwei Tao, Lang Mei, Xiaohan Yu, Liyan Liu, Chong Chen, Wentao Zhang
cs.AI
摘要
多模態大型語言模型(MLLMs)在視覺理解與推理方面取得了顯著進展,但其靜態參數化知識限制了其解決知識密集型且動態變化的開放世界問題的能力。為突破此一限制,多模態深度搜尋已成為開放世界資訊獲取的關鍵方向,並從單輪事實檢索演進至由視覺證據引導的長視野、多輪搜尋。然而,現有方法通常將視覺僅限制在輸入或答案階段,忽略了其在中間推理中的作用,且缺乏針對長視野互動的專門設計。因此,視覺證據難以驅動持續性檢索,制約了互動深度與推理範圍。為解決上述限制,我們提出了DeepVoyager-VL,一個用於視覺在環搜尋的長視野多模態深度搜尋框架。具體而言,我們構建多模態事件圖以驅動數據合成,生成具有中間視覺依賴與長推理鏈的問題。隨後,我們設計了具備主動視覺獲取與按需圖像加載能力的智能體框架。最後,我們在合成數據上對模型進行微調,無需強化學習。在十個多模態搜尋基準上的大量實驗驗證了我們方法的有效性。
English
Multimodal large language models (MLLMs) have advanced visual understanding and reasoning, yet their static parametric knowledge limits their ability to address knowledge-intensive and dynamically evolving open-world problems. To move beyond this limitation, multimodal deep search has emerged as a key direction for open-world information access, evolving from single-turn factual retrieval toward long-horizon, multi-turn search guided by visual evidence. However, existing methods typically confine vision to the input or answer stage, overlooking its role in intermediate reasoning, and lack designs tailored to long-horizon interaction. Consequently, visual evidence rarely drives continued retrieval, constraining both interaction depth and reasoning span. To address these limitations, we propose DeepVoyager-VL, a long-horizon multimodal deep-search framework for vision-in-the-loop search. Specifically, we construct a multimodal event graph to drive data synthesis, yielding problems with intermediate visual dependencies and long reasoning chains. We then design an agent framework for active visual acquisition and on-demand image loading. Finally, we fine-tune models on the synthesized data without reinforcement learning. Extensive experiments across ten multimodal search benchmarks demonstrate the effectiveness of our method.