V-Rubrics: 루브릭 기반 강화학습을 통한 시각적 충실성
V-Rubrics: Visual Faithfulness via Rubric-Based Reinforcement Learning
August 26, 2026
저자: Shulin Tian, Minglun Li, Yuhao Dong, Hao Ding, Jiarui Yao, Haiwen Diao, Jingkang Yang, Hongyuan Zhu, Ziwei Liu
cs.AI
초록
비전-언어 모델은 시각적 증거에 충분히 근거하지 않은 유창한 답변을 생성할 수 있다. 하나의 근거 없는 객체, 차트 값, 혹은 중간 추론이 다른 면에서는 그럴듯한 응답을 훼손할 수 있다. 우리는 이것이 멀티모달 후속 학습에서의 신용 할당 실패라고 주장한다. 스칼라 결과 보상은 답변이 수용 가능한지 여부를 나타낼 뿐, 어떤 시각적 사실이 근거를 갖는지, 어떤 추론 단계가 유효한지, 어떤 지시 제약이 지켜지지 않았는지를 식별하지 못한다. 우리는 시각적 루브릭 기반 강화 학습(Visual Rubrics-Based Reinforcement Learning)을 도입한다. 이는 참조 응답을 원자적 명제로 분해하고 생성된 답변을 시각적 충실성(VF), 추론 일관성(RC), 지시 수행(IF) 측면에서 채점한다. 결과로 얻어진 루브릭 항목들은 구조화된 부분 점수를 제공하며, 지지 증거 스팬이 사용 가능할 때 루브릭 신용을 국소화한다. 우리는 먼저 공개된 OpenMMReasoner-SFT-874K 코퍼스에서 Qwen3-VL-8B-Instruct를 미세 조정하고 OpenMMReasoner의 콜드 스타트 데이터 레시피를 적용하여 SFT 체크포인트를 확보한다. 우리는 17개의 시각적 근거 소스에서 50,248개 예제로 구성된 학습 세트인 V-Rubrics 50K를 구축하되, 규칙 기반 필터를 먼저 적용하고 거부 샘플링 점수에서 예제 난이도를 도출한 다음, 동일한 구조화된 프롬프트와 프로토콜에 따라 모든 예제를 Gemini-3-Pro로 주석하였다. 우리는 동일한 SFT 체크포인트를 기반으로, 구성 요소별, 접두사 국소화된 루브릭 신용을 사용하여 모델을 학습한다. 실험 결과, 우리의 루브릭 기반 GRPO는 공유 SFT 기준선과 답변 전용 GRPO보다 성능이 향상되었으며, 지식 중심 및 시각적 근거 추론 벤치마크에서 가장 큰 개선을 보였다. 이 결과는 루브릭이 시각적 후속 학습을 위한 유용한 보상 추상화임을 보여준다.
English
Vision-language models can produce fluent answers that are insufficiently grounded in the visual evidence: a single unsupported object, chart value, or intermediate inference can undermine an otherwise plausible response. We argue that this is a credit-assignment failure in multimodal post-training. Scalar outcome rewards indicate whether an answer is acceptable, but do not identify which visual facts are grounded, which reasoning steps are valid, or which instruction constraints are missed. We introduce Visual Rubrics-Based Reinforcement Learning, which decomposes reference responses into atomic propositions and scores generated answers along Visual Faithfulness (VF), Reasoning Consistency (RC), and Instruction Following (IF). The resulting rubric items provide structured partial credit and localize rubric credit when supporting evidence spans are available. We first obtain an SFT checkpoint by fine-tuning Qwen3-VL-8B-Instruct on the public OpenMMReasoner-SFT-874K corpus, adapting OpenMMReasoner's cold-start data recipe. We construct V-Rubrics 50K, a 50,248-example training set from 17 visually grounded sources, by applying rule-based filters before deriving example difficulty from rejection-sampling scores and then annotating every example with Gemini-3-Pro under the same structured prompt and protocol. We train our model based on the same SFT checkpoint using component-wise, prefix-localized rubric credit. Experiments show that our rubricbased GRPO improves over both the shared SFT baseline and answer-only GRPO, with the largest gains on knowledge-oriented and visually grounded reasoning benchmarks. The results show rubrics as a useful reward abstraction for visual post-training.