ChatPaper.aiChatPaper

VibeWorlding: 멀티모달 에이전트가 3D 오픈 월드를 엔드투엔드로 구축할 수 있는가?

VibeWorlding: Can Multimodal Agents Construct 3D Open Worlds End-to-End?

August 15, 2026
저자: Yansong Ning, Jingwen Ye, Zhongkai Wu, Yang Sun, Yiqin Zhu, Xingyi Li, Weidong Zhang, Hao Liu
cs.AI

초록

사용자 질의로부터 인터랙티브 3D 오픈 월드를 구축하는 것은 중요하다. 그러나 기존 방법들은 주로 이상화되고 단순한 질의에 대해 평가되어 왔으며, 이로 인해 멀티모달 에이전트가 사용자 의도를 이해하고, 3D 도구를 사용하며, 텍스트 및 시각적 3D 세계 정보에 대해 추론하는 능력을 체계적으로 분석하고 비교하기 어렵다. 이러한 문제를 해결하기 위해 우리는 VibeWorlding을 제안한다. 이는 바이브 월딩 에이전트를 벤치마킹하고 훈련하기 위한 통합 프레임워크로, 바이브 월딩 에이전트는 사용자 의도를 자율적으로 추론하고, 장면 레이아웃을 계획하며, 3D 도구를 호출하고, 다중 턴 에이전트-환경 상호작용 과정에서 멀티모달 피드백을 반영할 수 있는 멀티모달 에이전트이다. 이를 위해 우리는 먼저 VWE-BENCH를 구축한다. 이 벤치마크는 2,616개의 고품질 3D 자산, 323개의 인간 주석 시드 3D 월드, 6,828개의 역합성 멀티모달 사용자 질의로 구성되며, 정답이 포함된 검증된 질의와 정교하게 설계된 루브릭을 갖춘 미검증 질의로 분할된다. 또한, 우리는 VibeWorlding-Gym을 개발한다. 이는 다음을 통합하는 공동 멀티모달 강화 학습 사후 훈련 프레임워크이다: (1) 자산 검색, 편집, 이미지 렌더링을 MCP 도구로 통합하는 샌드박스 환경, (2) 물리적 실현 가능성 검증과 의도 충족 검증을 결합한 루브릭 기반 검증기로, 공정한 모델 평가와 확장 가능한 멀티모달 강화 학습 보상 서비스를 모두 지원한다. 우리의 실험은 현재 최첨단 멀티모달 대규모 언어 모델들이 바이브 월딩 에이전트 작업을 해결하는 데 여전히 부족하며, GPT-5.5와 Qwen3.8-Max조차도 60% 미만의 성공률을 기록함을 보여준다. 또한, 우리는 그 병목 지점이 정밀한 3D 세계 편집에 있음을 추적한다. 나아가 강화 학습 훈련이 이러한 약점을 완화할 수 있으며, 오픈소스 멀티모달 대규모 언어 모델이 폐쇄형 최첨단 모델을 능가할 수 있게 한다는 것을 발견한다. 우리의 VibeWorlder-8B는 최첨단 멀티모달 대규모 언어 모델과 필적하는 성능을 보이며, 주력 모델인 VibeWorlder-30B-A3B는 평가된 모든 모델 중 최고의 전반적 Pass@1을 달성한다.
English
Constructing an interactive 3D open world from a user query is important. However, existing methods are primarily evaluated on idealized, simple queries, making it difficult to systematically analyze and compare how multimodal agents understand user intent, use 3D tools, and reason over textual and visual 3D world information. To this end, we propose VibeWorlding, a unified framework for benchmarking and training vibe worlding agents: a multimodal agent that can autonomously infer user intent, plan scene layout, invoke 3D tools, and reflect on the multimodal feedback in a multi-turn agent-environment interaction process. To achieve this, we first build VWE-BENCH, a benchmark of 2,616 high-quality 3D assets, 323 human-annotated seed 3D worlds, and 6,828 reverse-synthesized multimodal user queries, split into verified queries with ground-truth and unverified queries with carefully designed rubrics. Moreover, we develop VibeWorlding-Gym, a joint multimodal RL post-training framework that integrates (1) a sandbox environment unifying asset retrieval, editing, and image rendering as MCP tools, and (2) a rubric-based verifier that combines physical feasibility and intent fulfillment verification, supporting both fair model evaluation and scalable multimodal RL reward service. Our experiments show that current frontier MLLMs are far from solving the vibe worlding agent task, with even GPT-5.5 and Qwen3.8-Max reaching below 60% success rate, and trace the bottleneck to precise 3D world editing. We further find that RL training can ease this weakness and enable open-source MLLMs to even surpass closed-source frontiers: our VibeWorlder-8B is comparable to frontier MLLMs, while our flagship VibeWorlder-30B-A3B attains the best overall Pass@1 among all evaluated models.