ChatPaper.aiChatPaper

DeepVoyager-VL:激励长时程多模态智能体的视觉在环搜索

DeepVoyager-VL: Incentivizing Vision-in-the-Loop Search for Long-Horizon Multimodal Agents

August 3, 2026
作者: Huanyao Zhang, Jiepeng Zhou, Runhao Zhao, Yanzhe Shan, Jiaoyang Chen, Bowen Zhou, Bo Li, Fang Wang, Jialong Wu, Zhengwei Tao, Lang Mei, Xiaohan Yu, Liyan Liu, Chong Chen, Wentao Zhang
cs.AI

摘要

多模态大语言模型(MLLMs)在视觉理解与推理方面取得了显著进展,但其静态参数化知识限制了其应对知识密集型及动态演变的开放世界问题的能力。为突破这一局限,多模态深度搜索应运而生,成为开放世界信息获取的关键方向,并正从单轮事实检索向由视觉证据引导的长程多轮搜索演进。然而,现有方法通常将视觉能力局限于输入或答案生成阶段,忽视了其在中间推理过程中的作用,且缺乏面向长程交互的专门设计。因此,视觉证据很少能驱动持续检索,交互深度与推理跨度均受到制约。为解决上述问题,我们提出了DeepVoyager-VL,一种面向视觉在环搜索的长程多模态深度搜索框架。具体而言,我们构建多模态事件图以驱动数据合成,生成具有中间视觉依赖关系和长推理链的问题;随后设计用于主动视觉获取与按需图像加载的智能体框架;最后在合成数据上对模型进行微调,无需强化学习。在十个多模态搜索基准上的大量实验验证了所提方法的有效性。
English
Multimodal large language models (MLLMs) have advanced visual understanding and reasoning, yet their static parametric knowledge limits their ability to address knowledge-intensive and dynamically evolving open-world problems. To move beyond this limitation, multimodal deep search has emerged as a key direction for open-world information access, evolving from single-turn factual retrieval toward long-horizon, multi-turn search guided by visual evidence. However, existing methods typically confine vision to the input or answer stage, overlooking its role in intermediate reasoning, and lack designs tailored to long-horizon interaction. Consequently, visual evidence rarely drives continued retrieval, constraining both interaction depth and reasoning span. To address these limitations, we propose DeepVoyager-VL, a long-horizon multimodal deep-search framework for vision-in-the-loop search. Specifically, we construct a multimodal event graph to drive data synthesis, yielding problems with intermediate visual dependencies and long reasoning chains. We then design an agent framework for active visual acquisition and on-demand image loading. Finally, we fine-tune models on the synthesized data without reinforcement learning. Extensive experiments across ten multimodal search benchmarks demonstrate the effectiveness of our method.