VBVR-Pro: 네이티브 시각 추론을 위한 확장 가능하고 검증 가능한 스위트
VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning
August 26, 2026
저자: Junxiang Xu, Ruisi Wang, Fanyi Pu, Maijunxian Wang, Ran Ji, Tongxi Zhou, Chenyang Gu, Jing Zuo, Hongcan Xiao, Yimeng Geng, Wanqi Yin, Wei Chen, Oscar Qian, Zhengan Yan, Ziqi Huang, Haiwen Diao, Liang Pan, Bo Li, Xiangyu Fan, Dezhi Luo, Fengyuan Yu, Zehong Zhao, Qingying Gao, Tinghui Zhu, Yilan Zhang, Jingqi Tong, Pinyuan Feng, Zhengze Jiang, Letian Wang, Ziyu Guo, Renrui Zhang, Jieneng Chen, Sonia Joseph, Constantin Venhoff, Saman Motamed, Mengyue Yang, Chandra Sripada, Alan Yuille, Philip Torr, Lvmin Zhang, Vikash Kumar, Daniel Khashabi, Nikolaus Kriegeskorte, Raphaël Millière, Vincent C. Müller, Anyi Rao, Quan Wang, Ziwei Liu, Dahua Lin, Lei Yang, Hokin Deng, Zhongang Cai
cs.AI
초록
시각 고유 추론(Native visual reasoning)은 시각 생성을 추론의 매개체 자체로 간주한다. 즉, 시각 상태(이미지와 비디오)는 단순히 이해해야 할 입력이나 렌더링할 출력이 아니라, 언어를 넘어선 문제 해결을 위한 일급 기반(first-class substrate)이다. 그러나 확장 가능한 훈련 작업, 신뢰할 수 있는 피드백, 그리고 생성 기반 간의 통제된 비교가 부족하여 진전이 병목에 막혀 있다. 본 연구에서는 생성을 통한 시각 고유 추론을 훈련, 검증, 최적화, 실험적 통제가 모두 가능하게 만드는 폐쇄 루프 테스트베드인 VBVR-Pro를 소개한다.
1) 작업 확장(Task scaling): VBVR-Pro는 시각 추론을 절차적으로 생성된 300개의 작업으로 이루어진 통제된 작업 공간으로 변환한다. VBVR-Pro로 훈련된 모델은 RISE-Video, MME-CoF-Pro, BabyVision 등 7개의 외부 시각 추론 벤치마크에서 제안된 작업 모음을 넘어 강력한 전이 성능을 보인다.
2) 검증 가능한 보상(Verifiable rewards): VBVR-Pro는 작업 기반 평가를 위한 검증 가능한 보상 스코어러를 제공한다. 주요 다중모달 대규모 언어 모델(MLLM)을 평가자로 활용한 체계적 연구를 통해, 널리 사용되는 VLM-as-a-judge 패러다임의 반복적 실패 양상을 식별한다. 이와 대조적으로, 제안된 스코어러는 결정론적이고 작업별 규칙에 기반하여 인간 판단과 세밀하게 정렬된다. 중요하게도, 이는 대규모 다중 작업 강화 학습을 위한 신뢰할 수 있는 보상 신호로 기능하며, 다양한 시각 추론 작업에서 강화 학습 이후 더 강력한 성능을 입증한다.
3) 메커니즘 연구(Mechanism study): VBVR-Pro는 30개 이상의 이미지, 비디오, 인터리브 생성기에 걸친 통제된 모달리티 연구를 가능하게 한다. 분석 결과, 지속적인 시공간 상태 추적이 필요한 작업에서는 비디오 생성이 가장 강력한 반면, 인터리브 생성은 계산 효율적인 대안을 제공한다. 중요하게도, 절제 연구(ablation)와 프로빙(probing)은 시각 추론에 핵심적인 시각 고유 궤적(vision-native trajectory)의 존재를 시사한다.
모든 데이터, 모델, 스코어러, 코드를 공개한다.
English
Native visual reasoning treats visual generation as the medium of reasoning itself: visual states (i.e. images and videos) are not merely inputs to be understood or outputs to be rendered, but first-class substrates for problem solving beyond language. Yet progress remains bottlenecked by the lack of scalable training tasks, reliable feedback, and controlled comparisons across generative substrates. In this work, we introduce VBVR-Pro, a closed-loop testbed that makes native visual reasoning through generation trainable, verifiable, optimizable, and experimentally controllable. 1) Task scaling. VBVR-Pro turns visual reasoning into a controlled task space of 300 procedurally generated tasks. Models trained on VBVR-Pro show strong transfer beyond the proposed suite across seven external visual reasoning benchmarks such as RISE-Video, MME-CoF-Pro, and BabyVision. 2) Verifiable rewards. VBVR-Pro provides verifiable reward scorers for task-grounded evaluation. Through a systematic study of leading MLLMs as judges, we identify recurring failure modes of the prevalent VLM-as-a-judge paradigm. In contrast, the proposed scorers are grounded in deterministic, task-specific rules, achieve fine-grained alignment with human judgments. Importantly, they serve as reliable reward signals for large-scale multi-task reinforcement learning and demonstrate stronger post-RL performance across visual reasoning tasks. 3) Mechanism study. VBVR-Pro enables controlled modality studies across more than 30 image, video, and interleaved generators. Our analysis shows that video generation remains strongest for tasks requiring persistent spatiotemporal state tracking, while interleaved generation provides a compute-efficient alternative. Critically, ablations and probing suggest the presence of vision-native trajectories that are crucial to visual reasoning. We release all data, models, scorers, and code.