ChatPaper.aiChatPaper

Evidence-RL: 証拠集約型の視覚的推論を目指して

Evidence-RL: Towards Evidence-intensive Visual Reasoning

August 8, 2026
著者: Haojie Huang, Xinlei Yu, Chengming Xu, Zhangquan Chen, Cheng Yang, Qingdong He, Yu Yang, Jiangning Zhang, Xiaobin Hu
cs.AI

要旨

視覚言語モデル(VLM)は、言語的な事前知識、データセットのショートカット、または無関係な視覚的文脈ではなく、具体的な画像エビデンスに基づいて回答すべきである。既存の知覚を考慮したポストトレーニング手法は、グローバルな摂動やアテンションによる代理指標を通じて画像の利用を促進するが、サンプリングされた回答がそれを裏付ける局所的なエビデンスに因果的に依存しているかどうかは検証しない。本稿では、VLMのグラウンディングに対する学習時エビデンス監査であるCounterfactual Evidence Disentanglement(CED)を提案する。CEDは各応答に対して、オブジェクト中心のエビデンス領域を中和し、得られたサポート低下を対応する非エビデンス領域と比較する。我々はこのシグナルをGRPO内で回答の正しさと組み合わせ、ショートカット経路やノイズ経路ではなくエビデンス経路に依存する正解回答を報酬づける。CEDは弱いオブジェクトレベルの候補提案を使用し、質問固有のエビデンスアノテーションを必要とせず、推論時のオーバーヘッドも追加しない。9つの公開ベンチマークと4つのバックボーンにわたって、CEDは既存の強化学習ベースのポストトレーニング手法を上回り、対象を絞った分析によりそのオブジェクト中心のシグナルが検証されている。
English
Vision-Language Models (VLMs) should answer from concrete image evidence rather than language priors, dataset shortcuts, or irrelevant visual context. Existing perception-aware post-training methods encourage image use through global perturbations or attention proxies, but they do not test whether a sampled answer causally depends on the local evidence that supports it. We propose Counterfactual Evidence Disentanglement (CED), a training-time evidence audit for VLM grounding. For each response, CED neutralizes an object-centric Evidence Region and compares the resulting support drop against matched non-evidence Regions. We combine this signal with answer correctness inside GRPO, rewarding correct answers that rely on the evidence path rather than shortcut or nuisance paths. CED uses weak object-level proposals, requires no question-specific evidence annotations, and adds no inference-time overhead. Across nine public benchmarks and four backbones, CED outperforms prior RL-based post-training methods, with targeted analyses verifying its object-centric signal.