ChatPaper.aiChatPaper

VBVR-Pro:一個可擴展且可驗證的原生視覺推理套件

VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning

August 26, 2026
作者: Junxiang Xu, Ruisi Wang, Fanyi Pu, Maijunxian Wang, Ran Ji, Tongxi Zhou, Chenyang Gu, Jing Zuo, Hongcan Xiao, Yimeng Geng, Wanqi Yin, Wei Chen, Oscar Qian, Zhengan Yan, Ziqi Huang, Haiwen Diao, Liang Pan, Bo Li, Xiangyu Fan, Dezhi Luo, Fengyuan Yu, Zehong Zhao, Qingying Gao, Tinghui Zhu, Yilan Zhang, Jingqi Tong, Pinyuan Feng, Zhengze Jiang, Letian Wang, Ziyu Guo, Renrui Zhang, Jieneng Chen, Sonia Joseph, Constantin Venhoff, Saman Motamed, Mengyue Yang, Chandra Sripada, Alan Yuille, Philip Torr, Lvmin Zhang, Vikash Kumar, Daniel Khashabi, Nikolaus Kriegeskorte, Raphaël Millière, Vincent C. Müller, Anyi Rao, Quan Wang, Ziwei Liu, Dahua Lin, Lei Yang, Hokin Deng, Zhongang Cai
cs.AI

摘要

原生視覺推理將視覺生成視為推理本身的媒介:視覺狀態(即圖像和影片)不僅僅是被理解的輸入或被渲染的輸出,而是超越語言的問題求解的一等基底。然而,由於缺乏可擴展的訓練任務、可靠的回饋,以及跨生成基底的受控比較,進展仍然受限。在本工作中,我們提出 VBVR-Pro,一個閉環測試平台,使以生成作為媒介的原生視覺推理變得可訓練、可驗證、可最佳化且可實驗控制。1) 任務擴展。VBVR-Pro 將視覺推理轉化為包含 300 個程序化生成任務的受控任務空間。在 VBVR-Pro 上訓練的模型,在 RISE-Video、MME-CoF-Pro 和 BabyVision 等七個外部視覺推理基準上,展現出超越本套件的強大遷移能力。2) 可驗證獎勵。VBVR-Pro 為基於任務的評估提供了可驗證的獎勵評分器。透過對領先多模態大語言模型作為評判者的系統性研究,我們識別出當前流行的「視覺語言模型作為評判者」範式的反覆出現的失敗模式。相比之下,所提出的評分器植基於確定性、任務特定的規則,並實現了與人類判斷的細粒度對齊。重要的是,它們可作為大規模多任務強化學習的可靠獎勵訊號,並在各種視覺推理任務上展現出更強的強化學習後性能。3) 機制研究。VBVR-Pro 支援對超過 30 種圖像、影片和交錯生成器進行受控模態研究。我們的分析顯示,對於需要持續時空狀態追蹤的任務,影片生成仍然是最強的,而交錯生成則提供了一種計算效率高的替代方案。至關重要的是,消融實驗和探測表明存在對視覺推理至關重要的視覺原生軌跡。我們發布所有資料、模型、評分器和程式碼。
English
Native visual reasoning treats visual generation as the medium of reasoning itself: visual states (i.e. images and videos) are not merely inputs to be understood or outputs to be rendered, but first-class substrates for problem solving beyond language. Yet progress remains bottlenecked by the lack of scalable training tasks, reliable feedback, and controlled comparisons across generative substrates. In this work, we introduce VBVR-Pro, a closed-loop testbed that makes native visual reasoning through generation trainable, verifiable, optimizable, and experimentally controllable. 1) Task scaling. VBVR-Pro turns visual reasoning into a controlled task space of 300 procedurally generated tasks. Models trained on VBVR-Pro show strong transfer beyond the proposed suite across seven external visual reasoning benchmarks such as RISE-Video, MME-CoF-Pro, and BabyVision. 2) Verifiable rewards. VBVR-Pro provides verifiable reward scorers for task-grounded evaluation. Through a systematic study of leading MLLMs as judges, we identify recurring failure modes of the prevalent VLM-as-a-judge paradigm. In contrast, the proposed scorers are grounded in deterministic, task-specific rules, achieve fine-grained alignment with human judgments. Importantly, they serve as reliable reward signals for large-scale multi-task reinforcement learning and demonstrate stronger post-RL performance across visual reasoning tasks. 3) Mechanism study. VBVR-Pro enables controlled modality studies across more than 30 image, video, and interleaved generators. Our analysis shows that video generation remains strongest for tasks requiring persistent spatiotemporal state tracking, while interleaved generation provides a compute-efficient alternative. Critically, ablations and probing suggest the presence of vision-native trajectories that are crucial to visual reasoning. We release all data, models, scorers, and code.