Evidence-RL: 증거 집약적 시각적 추론을 향하여
Evidence-RL: Towards Evidence-intensive Visual Reasoning
August 8, 2026
저자: Haojie Huang, Xinlei Yu, Chengming Xu, Zhangquan Chen, Cheng Yang, Qingdong He, Yu Yang, Jiangning Zhang, Xiaobin Hu
cs.AI
초록
비전-언어 모델(VLM)은 언어 사전 지식, 데이터셋 지름길, 또는 무관한 시각적 맥락보다는 구체적인 이미지 증거를 바탕으로 응답해야 한다. 기존의 지각 기반 사후 학습 방법들은 전역 섭동이나 주의 프록시를 통해 이미지 사용을 장려하지만, 샘플링된 응답이 이를 뒷받침하는 국소적 증거에 인과적으로 의존하는지 여부는 검증하지 않는다. 본 논문에서는 VLM 근거 설정을 위한 학습 시점 증거 감사 방법인 반사실적 증거 분리(CED)를 제안한다. CED는 각 응답에 대해 객체 중심 증거 영역을 중화시키고, 이로 인한 지지도 하락을 매칭된 비증거 영역과 비교한다. 우리는 이 신호를 GRPO 내부의 답변 정확도와 결합하여, 지름길 경로나 교란 경로가 아닌 증거 경로에 의존하는 정답에 보상을 부여한다. CED는 약한 객체 수준 제안을 활용하며, 질문별 증거 주석이 필요 없고, 추론 시 오버헤드도 추가하지 않는다. 아홉 개의 공개 벤치마크와 네 개의 백본에 걸친 실험에서, CED는 기존의 강화학습 기반 사후 학습 방법들을 능가하며, 정밀 분석을 통해 객체 중심 신호의 유효성을 검증한다.
English
Vision-Language Models (VLMs) should answer from concrete image evidence rather than language priors, dataset shortcuts, or irrelevant visual context. Existing perception-aware post-training methods encourage image use through global perturbations or attention proxies, but they do not test whether a sampled answer causally depends on the local evidence that supports it. We propose Counterfactual Evidence Disentanglement (CED), a training-time evidence audit for VLM grounding. For each response, CED neutralizes an object-centric Evidence Region and compares the resulting support drop against matched non-evidence Regions. We combine this signal with answer correctness inside GRPO, rewarding correct answers that rely on the evidence path rather than shortcut or nuisance paths. CED uses weak object-level proposals, requires no question-specific evidence annotations, and adds no inference-time overhead. Across nine public benchmarks and four backbones, CED outperforms prior RL-based post-training methods, with targeted analyses verifying its object-centric signal.