VBVR-Pro: ネイティブ視覚推論のためのスケーラブルかつ検証可能なスイート
VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning
August 26, 2026
著者: Junxiang Xu, Ruisi Wang, Fanyi Pu, Maijunxian Wang, Ran Ji, Tongxi Zhou, Chenyang Gu, Jing Zuo, Hongcan Xiao, Yimeng Geng, Wanqi Yin, Wei Chen, Oscar Qian, Zhengan Yan, Ziqi Huang, Haiwen Diao, Liang Pan, Bo Li, Xiangyu Fan, Dezhi Luo, Fengyuan Yu, Zehong Zhao, Qingying Gao, Tinghui Zhu, Yilan Zhang, Jingqi Tong, Pinyuan Feng, Zhengze Jiang, Letian Wang, Ziyu Guo, Renrui Zhang, Jieneng Chen, Sonia Joseph, Constantin Venhoff, Saman Motamed, Mengyue Yang, Chandra Sripada, Alan Yuille, Philip Torr, Lvmin Zhang, Vikash Kumar, Daniel Khashabi, Nikolaus Kriegeskorte, Raphaël Millière, Vincent C. Müller, Anyi Rao, Quan Wang, Ziwei Liu, Dahua Lin, Lei Yang, Hokin Deng, Zhongang Cai
cs.AI
要旨
ネイティブな視覚推論は、視覚生成を推論自体の媒体として扱う。すなわち、視覚状態(画像や動画)は単に理解すべき入力や描画すべき出力ではなく、言語を超えた問題解決のための第一級の基盤である。しかし、その進展は、スケーラブルな訓練タスク、信頼性の高いフィードバック、そして生成基盤間の制御された比較の欠如によって依然としてボトルネックに直面している。本研究では、生成を通じたネイティブな視覚推論を訓練可能・検証可能・最適化可能・実験的に制御可能にするクローズドループテストベッドVBVR-Proを紹介する。1) タスクのスケーリング。VBVR-Proは視覚推論を、手続き的に生成された300のタスクからなる制御されたタスク空間へと変換する。VBVR-Proで訓練されたモデルは、提案スイートを超えて、RISE-Video、MME-CoF-Pro、BabyVisionなどの7つの外部視覚推論ベンチマークにおいて強い転移を示す。2) 検証可能な報酬。VBVR-Proはタスクに基づいた評価のための検証可能な報酬スコアラーを提供する。主要なMLLMを判定者として用いる体系的な研究を通じて、普及しているVLM-as-a-judgeパラダイムの再発する障害モードを特定する。対照的に、提案するスコアラーは決定的でタスク固有のルールに基づいており、人間の判断との細かな整合性を達成する。重要なことに、これらは大規模マルチタスク強化学習のための信頼性の高い報酬信号として機能し、視覚推論タスク全体において強化学習後のより強い性能を実証する。3) メカニズム研究。VBVR-Proは30以上の画像・動画・インターリーブ生成器にわたる制御されたモダリティ研究を可能にする。我々の分析は、持続的な時空間状態追跡を必要とするタスクでは動画生成が最も強力であり、一方インターリーブ生成は計算効率の高い代替手段を提供することを示している。重要なことに、アブレーションとプロービングは、視覚推論に不可欠な視覚ネイティブな軌跡の存在を示唆している。我々はすべてのデータ、モデル、スコアラー、コードを公開する。
English
Native visual reasoning treats visual generation as the medium of reasoning itself: visual states (i.e. images and videos) are not merely inputs to be understood or outputs to be rendered, but first-class substrates for problem solving beyond language. Yet progress remains bottlenecked by the lack of scalable training tasks, reliable feedback, and controlled comparisons across generative substrates. In this work, we introduce VBVR-Pro, a closed-loop testbed that makes native visual reasoning through generation trainable, verifiable, optimizable, and experimentally controllable. 1) Task scaling. VBVR-Pro turns visual reasoning into a controlled task space of 300 procedurally generated tasks. Models trained on VBVR-Pro show strong transfer beyond the proposed suite across seven external visual reasoning benchmarks such as RISE-Video, MME-CoF-Pro, and BabyVision. 2) Verifiable rewards. VBVR-Pro provides verifiable reward scorers for task-grounded evaluation. Through a systematic study of leading MLLMs as judges, we identify recurring failure modes of the prevalent VLM-as-a-judge paradigm. In contrast, the proposed scorers are grounded in deterministic, task-specific rules, achieve fine-grained alignment with human judgments. Importantly, they serve as reliable reward signals for large-scale multi-task reinforcement learning and demonstrate stronger post-RL performance across visual reasoning tasks. 3) Mechanism study. VBVR-Pro enables controlled modality studies across more than 30 image, video, and interleaved generators. Our analysis shows that video generation remains strongest for tasks requiring persistent spatiotemporal state tracking, while interleaved generation provides a compute-efficient alternative. Critically, ablations and probing suggest the presence of vision-native trajectories that are crucial to visual reasoning. We release all data, models, scorers, and code.