ChatPaper.aiChatPaper

UniVR:视觉空间思维下的统一视觉推理

UniVR: Thinking in Visual Space for Unified Visual Reasoning

July 14, 2026
作者: Zhongwei Ren, Yunchao Wei, Yao Zhao, Weibo Gong, Xiao Liu, Anran Wang, Xiangtai Li, Xiaojie Jin
cs.AI

摘要

直接从原始视觉数据中学习广泛的世界知识是智能的基本能力。我们提出UniVR,这是首个研究如何从纯视觉演示中同时学习复杂推理、细粒度物理动力学和长期规划的探索。其核心在于VR-GRPO,一种结合全局级与步骤级互补奖励的强化学习范式。该方法在整个推理过程中强制保持逻辑连贯性与物理一致性,无需依赖特定任务启发式规则或图像-文本对。为训练和评估UniVR,我们构建了VR-X,一个涵盖长期操作、空间谜题和物理推理等16种多样化来源的大规模基准数据集。这是首个在纯视觉协议下评估这些异质能力的综合性测试集。值得注意的是,UniVR在VR-X上实现了高达25%的性能提升,其卓越的视觉推理能力也显著提升了多项多模态理解基准的表现。这些发现揭示了在视觉空间中进行推理的巨大潜力,所有代码、数据和模型均已开源以供进一步研究。
English
Learning broad world knowledge directly from raw visual data is a fundamental capability of intelligence. We introduce UniVR, the first investigation into simultaneously learning complex reasoning, fine-grained physical dynamics, and long-term planning from pure visual demonstrations. At its core, UniVR features VR-GRPO, a reinforcement learning paradigm with complementary global and step-level rewards. This approach enforces logical coherence and physical consistency throughout the reasoning process without requiring task-specific heuristics or image-text pairs. To train and evaluate UniVR, we construct VR-X, a large-scale benchmark curated from 16 diverse sources spanning long-horizon manipulation, spatial puzzles, and physical reasoning. It is the first comprehensive suite to assess these heterogeneous capabilities under a purely visual protocol. Remarkably, UniVR achieves up to a 25% improvement on VR-X, and its superior visual reasoning also boosts performance on various multimodal understanding benchmarks. These findings underscore the vast potential of reasoning within visual spaces, with all code, data, and models are open-sourced for further research.