UniVR: 視覚空間における思考による統合的視覚推論
UniVR: Thinking in Visual Space for Unified Visual Reasoning
July 14, 2026
著者: Zhongwei Ren, Yunchao Wei, Yao Zhao, Weibo Gong, Xiao Liu, Anran Wang, Xiangtai Li, Xiaojie Jin
cs.AI
要旨
生の視覚データから直接、広範な世界知識を学習することは、知能の基本的な能力である。本稿では、純粋な視覚的デモンストレーションから、複雑な推論、精細な物理ダイナミクス、長期計画を同時に学習する初の試みであるUniVRを紹介する。その中核となるのは、VR-GRPOと呼ばれる、相補的な全体報酬とステップ単位の報酬を備えた強化学習パラダイムである。この手法は、タスク固有のヒューリスティックスや画像-テキストペアを必要とせずに、推論プロセス全体を通じて論理的一貫性と物理的整合性を強制する。UniVRの学習と評価のために、長期操作、空間パズル、物理的推論にわたる16の多様なソースから厳選された大規模ベンチマークであるVR-Xを構築した。これは、純粋な視覚的プロトコルの下でこれらの異種能力を評価する初の包括的スイートである。注目すべきことに、UniVRはVR-X上で最大25%の改善を達成し、その優れた視覚的推論は様々なマルチモーダル理解ベンチマークの性能も向上させる。これらの知見は、視覚空間内での推論の広大な可能性を強調するものであり、すべてのコード、データ、モデルはさらなる研究のためにオープンソース化されている。
English
Learning broad world knowledge directly from raw visual data is a fundamental capability of intelligence. We introduce UniVR, the first investigation into simultaneously learning complex reasoning, fine-grained physical dynamics, and long-term planning from pure visual demonstrations. At its core, UniVR features VR-GRPO, a reinforcement learning paradigm with complementary global and step-level rewards. This approach enforces logical coherence and physical consistency throughout the reasoning process without requiring task-specific heuristics or image-text pairs. To train and evaluate UniVR, we construct VR-X, a large-scale benchmark curated from 16 diverse sources spanning long-horizon manipulation, spatial puzzles, and physical reasoning. It is the first comprehensive suite to assess these heterogeneous capabilities under a purely visual protocol. Remarkably, UniVR achieves up to a 25% improvement on VR-X, and its superior visual reasoning also boosts performance on various multimodal understanding benchmarks. These findings underscore the vast potential of reasoning within visual spaces, with all code, data, and models are open-sourced for further research.