DeepVoyager-VL:長期的マルチモーダルエージェントのためのビジョン・イン・ザ・ループ探索の動機付け
DeepVoyager-VL: Incentivizing Vision-in-the-Loop Search for Long-Horizon Multimodal Agents
August 3, 2026
著者: Huanyao Zhang, Jiepeng Zhou, Runhao Zhao, Yanzhe Shan, Jiaoyang Chen, Bowen Zhou, Bo Li, Fang Wang, Jialong Wu, Zhengwei Tao, Lang Mei, Xiaohan Yu, Liyan Liu, Chong Chen, Wentao Zhang
cs.AI
要旨
マルチモーダル大規模言語モデル(MLLMs)は視覚的理解と推論を大きく進歩させてきたが、その静的なパラメトリック知識は、知識集約的で動的に変化するオープンワールド問題への対処能力を制限している。この限界を克服するため、マルチモーダル深層検索がオープンワールド情報アクセスの重要方向として台頭し、単一ターンの事実検索から、視覚的証拠に導かれる長期的・多ターン検索へと進化してきた。しかしながら、既存手法は通常、視覚を入力段階または回答段階に限定しており、中間推論におけるその役割を見落とし、長期的な対話に適した設計も欠いている。その結果、視覚的証拠が継続的な検索を駆動することはほとんどなく、対話の深さと推論の広がりの両方が制約されている。これらの限界に対処するため、我々はビジョン・イン・ザ・ループ検索のための長期的マルチモーダル深層検索フレームワークであるDeepVoyager-VLを提案する。具体的には、マルチモーダルイベントグラフを構築してデータ合成を駆動し、中間の視覚的依存関係と長い推論チェーンを持つ問題を生成する。次に、能動的な視覚取得とオンデマンド画像ロードのためのエージェントフレームワークを設計する。最後に、強化学習を用いずに合成データ上でモデルをファインチューニングする。10のマルチモーダル検索ベンチマークにわたる広範な実験により、本手法の有効性が実証された。
English
Multimodal large language models (MLLMs) have advanced visual understanding and reasoning, yet their static parametric knowledge limits their ability to address knowledge-intensive and dynamically evolving open-world problems. To move beyond this limitation, multimodal deep search has emerged as a key direction for open-world information access, evolving from single-turn factual retrieval toward long-horizon, multi-turn search guided by visual evidence. However, existing methods typically confine vision to the input or answer stage, overlooking its role in intermediate reasoning, and lack designs tailored to long-horizon interaction. Consequently, visual evidence rarely drives continued retrieval, constraining both interaction depth and reasoning span. To address these limitations, we propose DeepVoyager-VL, a long-horizon multimodal deep-search framework for vision-in-the-loop search. Specifically, we construct a multimodal event graph to drive data synthesis, yielding problems with intermediate visual dependencies and long reasoning chains. We then design an agent framework for active visual acquisition and on-demand image loading. Finally, we fine-tune models on the synthesized data without reinforcement learning. Extensive experiments across ten multimodal search benchmarks demonstrate the effectiveness of our method.