VibeWorlding:マルチモーダルエージェントは3Dオープンワールドをエンドツーエンドで構築できるのか?
VibeWorlding: Can Multimodal Agents Construct 3D Open Worlds End-to-End?
August 15, 2026
著者: Yansong Ning, Jingwen Ye, Zhongkai Wu, Yang Sun, Yiqin Zhu, Xingyi Li, Weidong Zhang, Hao Liu
cs.AI
要旨
ユーザーのクエリからインタラクティブな3Dオープンワールドを構築することは重要である。しかし、既存手法は主に理想化された単純なクエリで評価されており、マルチモーダルエージェントがユーザーの意図を理解し、3Dツールを使用し、テキストおよび視覚的な3Dワールド情報に基づいて推論する方法を体系的に分析・比較することは困難である。この目的のため、我々はVibeWorldingを提案する。これは、vibe worldingエージェントをベンチマーク評価および訓練するための統一フレームワークであり、マルチターンのエージェント・環境相互作用プロセスにおいて、ユーザーの意図を自律的に推論し、シーンレイアウトを計画し、3Dツールを呼び出し、マルチモーダルフィードバックを反映することができるマルチモーダルエージェントである。これを実現するために、まずVWE-BENCHを構築する。これは、2,616点の高品質な3Dアセット、323件の人間がアノテーションしたシード3Dワールド、および6,828件の逆合成されたマルチモーダルユーザークエリからなるベンチマークであり、正解ラベル付きの検証済みクエリと、慎重に設計されたルーブリック付きの未検証クエリに分割される。さらに、我々はVibeWorlding-Gymを開発する。これは、アセット検索、編集、画像レンダリングをMCPツールとして統合するサンドボックス環境と、物理的実現可能性と意図充足の検証を組み合わせたルーブリックベースの検証器を統合した、マルチモーダルRL事後訓練のための統合フレームワークであり、公平なモデル評価とスケーラブルなマルチモーダルRL報酬サービスの両方をサポートする。我々の実験は、現在の最先端MLLMがvibe worldingエージェントタスクを解くには程遠く、GPT-5.5やQwen3.8-Maxでさえ成功率が60%未満であることを示しており、そのボトルネックは精密な3Dワールド編集にあると特定している。さらに、RL訓練がこの弱点を緩和し、オープンソースのMLLMがクローズドソースの最先端モデルを凌駕することさえ可能にすることを見出した。我々のVibeWorlder-8Bは最先端MLLMに匹敵し、一方、主力モデルであるVibeWorlder-30B-A3Bは、評価されたすべてのモデルの中で総合的に最高のPass@1を達成する。
English
Constructing an interactive 3D open world from a user query is important. However, existing methods are primarily evaluated on idealized, simple queries, making it difficult to systematically analyze and compare how multimodal agents understand user intent, use 3D tools, and reason over textual and visual 3D world information. To this end, we propose VibeWorlding, a unified framework for benchmarking and training vibe worlding agents: a multimodal agent that can autonomously infer user intent, plan scene layout, invoke 3D tools, and reflect on the multimodal feedback in a multi-turn agent-environment interaction process. To achieve this, we first build VWE-BENCH, a benchmark of 2,616 high-quality 3D assets, 323 human-annotated seed 3D worlds, and 6,828 reverse-synthesized multimodal user queries, split into verified queries with ground-truth and unverified queries with carefully designed rubrics. Moreover, we develop VibeWorlding-Gym, a joint multimodal RL post-training framework that integrates (1) a sandbox environment unifying asset retrieval, editing, and image rendering as MCP tools, and (2) a rubric-based verifier that combines physical feasibility and intent fulfillment verification, supporting both fair model evaluation and scalable multimodal RL reward service. Our experiments show that current frontier MLLMs are far from solving the vibe worlding agent task, with even GPT-5.5 and Qwen3.8-Max reaching below 60% success rate, and trace the bottleneck to precise 3D world editing. We further find that RL training can ease this weakness and enable open-source MLLMs to even surpass closed-source frontiers: our VibeWorlder-8B is comparable to frontier MLLMs, while our flagship VibeWorlder-30B-A3B attains the best overall Pass@1 among all evaluated models.