UniVR: 통합 시각 추론을 위한 시각 공간에서의 사고
UniVR: Thinking in Visual Space for Unified Visual Reasoning
July 14, 2026
저자: Zhongwei Ren, Yunchao Wei, Yao Zhao, Weibo Gong, Xiao Liu, Anran Wang, Xiangtai Li, Xiaojie Jin
cs.AI
초록
원시 시각 데이터로부터 직접 광범위한 세계 지식을 학습하는 것은 지능의 기본적인 능력이다. 본 연구에서는 순수 시각적 시연으로부터 복잡한 추론, 세부적인 물리적 동역학, 그리고 장기 계획을 동시에 학습하는 최초의 연구인 UniVR을 소개한다. UniVR의 핵심은 VR-GRPO로, 상호 보완적인 전역 및 단계별 보상을 갖춘 강화 학습 패러다임이다. 이 접근법은 과제별 휴리스틱이나 이미지-텍스트 쌍 없이도 추론 과정 전반에 걸쳐 논리적 일관성과 물리적 일관성을 강제한다. UniVR을 훈련하고 평가하기 위해, 장기 조작, 공간 퍼즐, 물리적 추론을 포괄하는 16개의 다양한 소스에서 선별된 대규모 벤치마크인 VR-X를 구축하였다. 이는 순수 시각적 프로토콜 하에서 이러한 이질적 능력들을 평가하는 최초의 종합적인 평가 세트이다. 주목할 만하게도, UniVR은 VR-X에서 최대 25%의 성능 향상을 달성하였으며, 그 뛰어난 시각적 추론 능력은 다양한 멀티모달 이해 벤치마크에서도 성능을 향상시킨다. 이러한 발견은 시각 공간 내에서의 추론에 대한 광범위한 잠재력을 강조하며, 모든 코드, 데이터, 모델은 추가 연구를 위해 오픈소스로 공개된다.
English
Learning broad world knowledge directly from raw visual data is a fundamental capability of intelligence. We introduce UniVR, the first investigation into simultaneously learning complex reasoning, fine-grained physical dynamics, and long-term planning from pure visual demonstrations. At its core, UniVR features VR-GRPO, a reinforcement learning paradigm with complementary global and step-level rewards. This approach enforces logical coherence and physical consistency throughout the reasoning process without requiring task-specific heuristics or image-text pairs. To train and evaluate UniVR, we construct VR-X, a large-scale benchmark curated from 16 diverse sources spanning long-horizon manipulation, spatial puzzles, and physical reasoning. It is the first comprehensive suite to assess these heterogeneous capabilities under a purely visual protocol. Remarkably, UniVR achieves up to a 25% improvement on VR-X, and its superior visual reasoning also boosts performance on various multimodal understanding benchmarks. These findings underscore the vast potential of reasoning within visual spaces, with all code, data, and models are open-sourced for further research.