VBVR-Pro:一种可扩展且可验证的原生视觉推理套件
VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning
August 26, 2026
作者: Junxiang Xu, Ruisi Wang, Fanyi Pu, Maijunxian Wang, Ran Ji, Tongxi Zhou, Chenyang Gu, Jing Zuo, Hongcan Xiao, Yimeng Geng, Wanqi Yin, Wei Chen, Oscar Qian, Zhengan Yan, Ziqi Huang, Haiwen Diao, Liang Pan, Bo Li, Xiangyu Fan, Dezhi Luo, Fengyuan Yu, Zehong Zhao, Qingying Gao, Tinghui Zhu, Yilan Zhang, Jingqi Tong, Pinyuan Feng, Zhengze Jiang, Letian Wang, Ziyu Guo, Renrui Zhang, Jieneng Chen, Sonia Joseph, Constantin Venhoff, Saman Motamed, Mengyue Yang, Chandra Sripada, Alan Yuille, Philip Torr, Lvmin Zhang, Vikash Kumar, Daniel Khashabi, Nikolaus Kriegeskorte, Raphaël Millière, Vincent C. Müller, Anyi Rao, Quan Wang, Ziwei Liu, Dahua Lin, Lei Yang, Hokin Deng, Zhongang Cai
cs.AI
摘要
原生视觉推理将视觉生成本身视为推理的媒介:视觉状态(即图像和视频)不仅是待理解的输入或待渲染的输出,更是超越语言的、用于问题求解的一等基板。然而,该领域的发展仍受制于可扩展训练任务、可靠反馈以及跨生成基板的受控比较的缺乏。在本工作中,我们提出VBVR-Pro,一个闭环测试平台,使得通过生成进行原生视觉推理变得可训练、可验证、可优化且实验可控。1)任务扩展。VBVR-Pro将视觉推理转化为一个包含300个程序化生成任务的受控任务空间。在VBVR-Pro上训练的模型在七个外部视觉推理基准(如RISE-Video、MME-CoF-Pro和BabyVision)上展现出超越所提套件的强迁移能力。2)可验证奖励。VBVR-Pro为基于任务的评估提供可验证的奖励评分器。通过对主流多模态大语言模型作为评判者的系统研究,我们识别出当前盛行的VLM-as-a-judge范式中反复出现的失败模式。相比之下,所提出的评分器基于确定性的、任务特定的规则,能够与人类判断实现细粒度对齐。重要的是,它们可作为大规模多任务强化学习的可靠奖励信号,并在视觉推理任务上展现出更强的强化学习后性能。3)机制研究。VBVR-Pro支持跨30余种图像、视频和交错生成器的受控模态研究。我们的分析表明,视频生成在需要持续时空状态追踪的任务中仍然最为强大,而交错生成提供了一种计算高效的替代方案。至关重要的是,消融实验与探测分析提示了视觉原生轨迹的存在,这些轨迹对视觉推理至关重要。我们发布全部数据、模型、评分器及代码。
English
Native visual reasoning treats visual generation as the medium of reasoning itself: visual states (i.e. images and videos) are not merely inputs to be understood or outputs to be rendered, but first-class substrates for problem solving beyond language. Yet progress remains bottlenecked by the lack of scalable training tasks, reliable feedback, and controlled comparisons across generative substrates. In this work, we introduce VBVR-Pro, a closed-loop testbed that makes native visual reasoning through generation trainable, verifiable, optimizable, and experimentally controllable. 1) Task scaling. VBVR-Pro turns visual reasoning into a controlled task space of 300 procedurally generated tasks. Models trained on VBVR-Pro show strong transfer beyond the proposed suite across seven external visual reasoning benchmarks such as RISE-Video, MME-CoF-Pro, and BabyVision. 2) Verifiable rewards. VBVR-Pro provides verifiable reward scorers for task-grounded evaluation. Through a systematic study of leading MLLMs as judges, we identify recurring failure modes of the prevalent VLM-as-a-judge paradigm. In contrast, the proposed scorers are grounded in deterministic, task-specific rules, achieve fine-grained alignment with human judgments. Importantly, they serve as reliable reward signals for large-scale multi-task reinforcement learning and demonstrate stronger post-RL performance across visual reasoning tasks. 3) Mechanism study. VBVR-Pro enables controlled modality studies across more than 30 image, video, and interleaved generators. Our analysis shows that video generation remains strongest for tasks requiring persistent spatiotemporal state tracking, while interleaved generation provides a compute-efficient alternative. Critically, ablations and probing suggest the presence of vision-native trajectories that are crucial to visual reasoning. We release all data, models, scorers, and code.