ChatPaper.aiChatPaper

Beacon:エージェント型視覚的推論をいつ、どのように実行するか

Beacon: Knowing When and How to Perform Agentic Visual Reasoning

July 30, 2026
著者: Qixun Wang, Yang Shi, Letian Cheng, Zhuoran Zhang, Yan He, Yuqi Tang, Qi Zhang, Xinlei Yu, Ruizhe Chen, Tianrun Xu, Yuanxing Zhang, Pengfei Wan, Haotian Wang, Xianghua Ying
cs.AI

要旨

エージェント型視覚推論の根本的な目的は、マルチモーダル大規模言語モデル(MLLMs)の複雑なタスクにおける成功率を向上させることであり、単に洗練されているものの非効率な推論パラダイムを備えさせることではない。本研究では、ツール使用の二つの重要な側面、すなわちモード適応性(MA)とツール効果(TE)を通じて、エージェント型視覚推論を再考する。モード適応性は、MLLMがツールが本当に必要であるタイミングを認識し、それに応じてツールを呼び出せるかどうかを特徴づけるものであり、不要な計算オーバーヘッドを回避しつつ、ツール支援を必要とする困難な問題における性能を向上させる。ツール効果はツール使用の実際の影響を特徴づける。すなわち、ツールはテキストのみの推論では解決できない問題に対してモデルの能力を拡張すべきであり、一方でツールなしでもモデルが既に解決できる問題に対しては追加のエラーを引き起こすことを回避すべきである。我々はこれら二つの特性を定量化するための包括的な分析を行い、既存のエージェント型視覚推論モデルはモード適応性が限定的であること、そして難しい例におけるツール使用による利得が、モデルが既に解ける簡単な例において生じる悪影響によって大きく相殺されることを実証的に明らかにする。これらの観察に動機づけられ、我々はより強力な全体的性能、改善されたモード適応性、そして真のツール由来の性能向上を達成する新しいエージェント型視覚推論モデルBeaconを提案する。Beaconの中核を成すのは、強化学習段階における「必要性認識型適応報酬(Necessity-Aware Adaptive Reward)」と「ヒント誘導型能力拡張(Hint-Guided Capability Expansion)」メカニズムであり、それぞれタスクの必要性に基づく適応的なツール呼び出しを促進し、最も困難な問題に対するモデルのツール使用能力を強化する。多様なベンチマークにわたる広範な実験により、Beaconの強力な全体的性能と、モード適応性およびツール効果の両方における大幅な改善が実証される。
English
The fundamental goal of agentic visual reasoning is to improve the success rate of multimodal large language models (MLLMs) on complex tasks, rather than merely equipping them with a sophisticated yet inefficient reasoning paradigm. In this work, we rethink agentic visual reasoning through two key dimensions of tool use: Mode Adaptiveness (MA) and Tool Effect (TE). Mode Adaptiveness characterizes whether an MLLM can recognize when tools are truly necessary and invoke them accordingly, thereby avoiding unnecessary computational overhead while improving performance on challenging problems that require tool assistance. Tool Effect characterizes the actual impact of tool use: tools should extend the model's capabilities on problems unsolvable through text-only reasoning, while avoiding additional errors on problems that the model can already solve without tools. We conduct a comprehensive analysis to quantify these two properties and empirically reveal that existing agentic visual reasoning models exhibit limited Mode Adaptiveness, while the gains produced by tool use on hard examples are largely offset by the harm introduced on easy examples that the models can already solve. Motivated by these observations, we propose Beacon, a novel agentic visual reasoning model that achieves stronger overall performance, improved Mode Adaptiveness, and genuine tool-induced performance gains. At the core of Beacon are the Necessity-Aware Adaptive Reward and the Hint-Guided Capability Expansion mechanism in the reinforcement learning stage, which respectively encourage adaptive tool invocation based on task necessity and strengthen the model's tool-use capability on the most challenging problems. Extensive experiments across diverse benchmarks demonstrate the strong overall performance of Beacon and its substantial improvements in both Mode Adaptiveness and Tool Effect.