VibeWorlding:多模態智能體能否端到端地建構3D開放世界?
VibeWorlding: Can Multimodal Agents Construct 3D Open Worlds End-to-End?
August 15, 2026
作者: Yansong Ning, Jingwen Ye, Zhongkai Wu, Yang Sun, Yiqin Zhu, Xingyi Li, Weidong Zhang, Hao Liu
cs.AI
摘要
根據用戶查詢構建互動式3D開放世界至關重要。然而,現有方法主要是在理想化、簡單的查詢上進行評估,使得難以系統性地分析和比較多模態智能體如何理解用戶意圖、使用3D工具,以及基於文本和視覺3D世界信息進行推理。為此,我們提出了VibeWorlding,一個用於基準測試和訓練vibe worlding智能體的統一框架:該智能體能夠在多輪智能體-環境交互過程中自主推斷用戶意圖、規劃場景佈局、調用3D工具,並反思多模態反饋。為實現此目標,我們首先構建了VWE-BENCH,一個包含2,616個高質量3D資產、323個人工標註的種子3D世界,以及6,828條逆向合成的多模態用戶查詢的基準,分為具有真實標註的已驗證查詢和帶有精心設計評分標準的未驗證查詢。此外,我們開發了VibeWorlding-Gym,一個聯合多模態強化學習後訓練框架,整合了(1)一個將資產檢索、編輯和圖像渲染統合為MCP工具的沙盒環境,以及(2)一個基於評分標準的驗證器,該驗證器結合物理可行性和意圖滿足驗證,同時支持公平的模型評估和可擴展的多模態強化學習獎勵服務。我們的實驗表明,當前前沿的多模態大語言模型(MLLM)遠未能解決vibe worlding智能體任務,即使是GPT-5.5和Qwen3.8-Max的成功率也低於60%;我們進一步將瓶頸追溯到精確的3D世界編輯。我們進一步發現,強化學習訓練可以緩解這一弱點,並使開源MLLM甚至能超越閉源前沿模型:我們的VibeWorlder-8B可與前沿MLLM相媲美,而我們的旗艦模型VibeWorlder-30B-A3B在所有評估模型中取得了最佳的整體Pass@1。
English
Constructing an interactive 3D open world from a user query is important. However, existing methods are primarily evaluated on idealized, simple queries, making it difficult to systematically analyze and compare how multimodal agents understand user intent, use 3D tools, and reason over textual and visual 3D world information. To this end, we propose VibeWorlding, a unified framework for benchmarking and training vibe worlding agents: a multimodal agent that can autonomously infer user intent, plan scene layout, invoke 3D tools, and reflect on the multimodal feedback in a multi-turn agent-environment interaction process. To achieve this, we first build VWE-BENCH, a benchmark of 2,616 high-quality 3D assets, 323 human-annotated seed 3D worlds, and 6,828 reverse-synthesized multimodal user queries, split into verified queries with ground-truth and unverified queries with carefully designed rubrics. Moreover, we develop VibeWorlding-Gym, a joint multimodal RL post-training framework that integrates (1) a sandbox environment unifying asset retrieval, editing, and image rendering as MCP tools, and (2) a rubric-based verifier that combines physical feasibility and intent fulfillment verification, supporting both fair model evaluation and scalable multimodal RL reward service. Our experiments show that current frontier MLLMs are far from solving the vibe worlding agent task, with even GPT-5.5 and Qwen3.8-Max reaching below 60% success rate, and trace the bottleneck to precise 3D world editing. We further find that RL training can ease this weakness and enable open-source MLLMs to even surpass closed-source frontiers: our VibeWorlder-8B is comparable to frontier MLLMs, while our flagship VibeWorlder-30B-A3B attains the best overall Pass@1 among all evaluated models.