UniVR:在視覺空間中思考以實現統一視覺推理
UniVR: Thinking in Visual Space for Unified Visual Reasoning
July 14, 2026
作者: Zhongwei Ren, Yunchao Wei, Yao Zhao, Weibo Gong, Xiao Liu, Anran Wang, Xiangtai Li, Xiaojie Jin
cs.AI
摘要
從原始視覺資料中學習廣泛的世界知識,是智能的一項基本能力。我們提出 UniVR,這是首度探討如何從純視覺示範中,同時學習複雜推理、細緻的物理動態以及長期規劃的研究。UniVR 的核心在於 VR-GRPO,這是一種結合全域與步驟層級獎勵的強化學習範式。此方法在推理過程中強化了邏輯一致性與物理一致性,無需依賴任務特定的啟發式規則或圖文配對。為了訓練與評估 UniVR,我們建構了 VR-X,這是一個大規模基準測試,彙整了來自 16 個不同來源的資料,涵蓋長期操作、空間拼圖與物理推理。這是首個在純視覺協議下評估這些異質能力的綜合性測試集。值得注意的是,UniVR 在 VR-X 上實現了高達 25% 的效能提升,而其卓越的視覺推理能力也提升了多項多模態理解基準測試的表現。這些發現凸顯了在視覺空間中進行推理的巨大潛力,所有程式碼、數據與模型均已開源,以供後續研究使用。
English
Learning broad world knowledge directly from raw visual data is a fundamental capability of intelligence. We introduce UniVR, the first investigation into simultaneously learning complex reasoning, fine-grained physical dynamics, and long-term planning from pure visual demonstrations. At its core, UniVR features VR-GRPO, a reinforcement learning paradigm with complementary global and step-level rewards. This approach enforces logical coherence and physical consistency throughout the reasoning process without requiring task-specific heuristics or image-text pairs. To train and evaluate UniVR, we construct VR-X, a large-scale benchmark curated from 16 diverse sources spanning long-horizon manipulation, spatial puzzles, and physical reasoning. It is the first comprehensive suite to assess these heterogeneous capabilities under a purely visual protocol. Remarkably, UniVR achieves up to a 25% improvement on VR-X, and its superior visual reasoning also boosts performance on various multimodal understanding benchmarks. These findings underscore the vast potential of reasoning within visual spaces, with all code, data, and models are open-sourced for further research.