DeepVoyager-VL: 장기 지평 멀티모달 에이전트를 위한 비전-인-더-루프 탐색 장려
DeepVoyager-VL: Incentivizing Vision-in-the-Loop Search for Long-Horizon Multimodal Agents
August 3, 2026
저자: Huanyao Zhang, Jiepeng Zhou, Runhao Zhao, Yanzhe Shan, Jiaoyang Chen, Bowen Zhou, Bo Li, Fang Wang, Jialong Wu, Zhengwei Tao, Lang Mei, Xiaohan Yu, Liyan Liu, Chong Chen, Wentao Zhang
cs.AI
초록
다중모달 대규모 언어 모델(MLLM)은 시각적 이해와 추론을 발전시켰지만, 정적 파라메트릭 지식으로 인해 지식 집약적이고 동적으로 변화하는 개방형 세계(open-world) 문제를 해결하는 데 한계가 있다. 이러한 한계를 극복하기 위해 다중모달 딥 서치(multimodal deep search)가 개방형 세계 정보 접근의 핵심 방향으로 부상했으며, 단일 턴 사실 검색에서 시각적 증거에 기반한 장기 지평(long-horizon) 다중 턴 검색으로 진화하고 있다. 그러나 기존 방법은 일반적으로 시각을 입력 또는 응답 단계에 국한시켜 중간 추론에서의 역할을 간과하며, 장기 지평 상호작용에 맞춘 설계가 부족하다. 그 결과, 시각적 증거가 지속적 검색을 주도하는 경우가 드물어 상호작용의 깊이와 추론 범위가 모두 제약된다. 이러한 한계를 해결하기 위해 우리는 시각-인-더-루프(vision-in-the-loop) 검색을 위한 장기 지평 다중모달 딥 서치 프레임워크인 DeepVoyager-VL을 제안한다. 구체적으로, 중간 시각적 의존성과 긴 추론 체인을 갖는 문제를 생성하기 위해 다중모달 이벤트 그래프를 구축하여 데이터 합성을 주도한다. 그다음 능동적 시각 획득과 온디맨드 이미지 로딩을 위한 에이전트 프레임워크를 설계한다. 마지막으로 강화 학습 없이 합성 데이터에 대해 모델을 파인튜닝한다. 10개의 다중모달 검색 벤치마크에 걸친 광범위한 실험을 통해 우리 방법의 효과성을 입증한다.
English
Multimodal large language models (MLLMs) have advanced visual understanding and reasoning, yet their static parametric knowledge limits their ability to address knowledge-intensive and dynamically evolving open-world problems. To move beyond this limitation, multimodal deep search has emerged as a key direction for open-world information access, evolving from single-turn factual retrieval toward long-horizon, multi-turn search guided by visual evidence. However, existing methods typically confine vision to the input or answer stage, overlooking its role in intermediate reasoning, and lack designs tailored to long-horizon interaction. Consequently, visual evidence rarely drives continued retrieval, constraining both interaction depth and reasoning span. To address these limitations, we propose DeepVoyager-VL, a long-horizon multimodal deep-search framework for vision-in-the-loop search. Specifically, we construct a multimodal event graph to drive data synthesis, yielding problems with intermediate visual dependencies and long reasoning chains. We then design an agent framework for active visual acquisition and on-demand image loading. Finally, we fine-tune models on the synthesized data without reinforcement learning. Extensive experiments across ten multimodal search benchmarks demonstrate the effectiveness of our method.