Beacon:知道何時以及如何進行代理式視覺推理
Beacon: Knowing When and How to Perform Agentic Visual Reasoning
July 30, 2026
作者: Qixun Wang, Yang Shi, Letian Cheng, Zhuoran Zhang, Yan He, Yuqi Tang, Qi Zhang, Xinlei Yu, Ruizhe Chen, Tianrun Xu, Yuanxing Zhang, Pengfei Wan, Haotian Wang, Xianghua Ying
cs.AI
摘要
智能體式視覺推理的根本目標,在於提升多模態大型語言模型(MLLMs)在複雜任務上的成功率,而非僅為其配備一套精緻卻低效的推理範式。在本工作中,我們透過工具使用的兩個關鍵維度重新審視智能體式視覺推理:模式適應性(MA)與工具效果(TE)。模式適應性用以表徵多模態大型語言模型能否辨識工具真正必要的時機並據此加以呼叫,藉此避免不必要的計算負擔,同時提升在需要工具輔助的困難問題上的表現。工具效果則用以表徵工具使用的實際影響:工具應在模型無法僅靠純文字推理解決的問題上擴展其能力,同時避免在模型無需工具即可解決的問題上引入額外錯誤。我們進行了全面的分析以量化這兩個性質,並實證揭示現有智能體式視覺推理模型的模式適應性有限,而工具使用在困難範例上帶來的增益,很大程度上被其在模型已能解決的簡單範例上所引入的損害所抵銷。受這些觀察啟發,我們提出了 Beacon,一個新穎的智能體式視覺推理模型,其在整體表現、模式適應性以及真正的工具誘發效能增益方面均取得更佳成果。Beacon 的核心在於強化學習階段的「必要性感知自適應獎勵」與「提示引導能力擴展」機制,前者分別根據任務必要性鼓勵自適應的工具呼叫,後者則強化模型在最具挑戰性問題上的工具使用能力。跨多樣化基準的廣泛實驗證明了 Beacon 的強大整體表現,以及其在模式適應性與工具效果兩方面的顯著提升。
English
The fundamental goal of agentic visual reasoning is to improve the success rate of multimodal large language models (MLLMs) on complex tasks, rather than merely equipping them with a sophisticated yet inefficient reasoning paradigm. In this work, we rethink agentic visual reasoning through two key dimensions of tool use: Mode Adaptiveness (MA) and Tool Effect (TE). Mode Adaptiveness characterizes whether an MLLM can recognize when tools are truly necessary and invoke them accordingly, thereby avoiding unnecessary computational overhead while improving performance on challenging problems that require tool assistance. Tool Effect characterizes the actual impact of tool use: tools should extend the model's capabilities on problems unsolvable through text-only reasoning, while avoiding additional errors on problems that the model can already solve without tools. We conduct a comprehensive analysis to quantify these two properties and empirically reveal that existing agentic visual reasoning models exhibit limited Mode Adaptiveness, while the gains produced by tool use on hard examples are largely offset by the harm introduced on easy examples that the models can already solve. Motivated by these observations, we propose Beacon, a novel agentic visual reasoning model that achieves stronger overall performance, improved Mode Adaptiveness, and genuine tool-induced performance gains. At the core of Beacon are the Necessity-Aware Adaptive Reward and the Hint-Guided Capability Expansion mechanism in the reinforcement learning stage, which respectively encourage adaptive tool invocation based on task necessity and strengthen the model's tool-use capability on the most challenging problems. Extensive experiments across diverse benchmarks demonstrate the strong overall performance of Beacon and its substantial improvements in both Mode Adaptiveness and Tool Effect.