ChatPaper.aiChatPaper

WorldClaw:規模化的智能體3D開放世界生成

WorldClaw: Agentic 3D Open-World Generation at Scale

August 5, 2026
作者: Chunchao Guo, Jinpeng Li, Yang Li, Zilong Huang
cs.AI

摘要

從開放式文本生成大規模、可自由探索的3D世界仍具挑戰性,因為系統必須同時維持全局空間一致性、豐富的局部內容,以及適合下游編輯和重用的顯式資產。我們提出 WorldClaw,這是一個完全智能體驅動、由粗到細的框架,用於開放世界3D場景生成。規劃智能體將文本提示轉換為區域、地形、資產、材質和空間關係的結構化規範。接著,WorldClaw 從語義佈局、可重用資產、生成式或程序化材質,以及區域感知的高度場,構建全局一致的地形基礎。對於細節要求高的區域,它生成地形條件化的組合,重建可編輯的紋理網格,並恢復它們在地形上的放置位置;基於渲染的智能體進一步細化地形、物體、外觀和接觸。在各種開放世界提示下,WorldClaw 生成具有一致空間組織、視覺上引人入勝的局部內容和可編輯實例級資產的大規模場景,同時保持一致的全局地形結構。
English
Generating large-scale, freely explorable 3D worlds from open-ended text remains challenging because a system must jointly maintain global spatial coherence, rich local content, and explicit assets suitable for downstream editing and reuse. We present WorldClaw, a fully agentic, coarse-to-fine framework for open-world 3D scene generation. Planning agents translate a text prompt into a structured specification of regions, terrain, assets, materials, and spatial relations. WorldClaw then builds a globally coherent terrain foundation from semantic layouts, reusable assets, generative or procedural materials, and a region-aware height field. For detail-demanding regions, it generates terrain-conditioned compositions, reconstructs editable textured meshes, and recovers their placement on the terrain; render-based agents further refine terrain, objects, appearance, and contacts. Across diverse open-world prompts, WorldClaw produces large-scale scenes with coherent spatial organization, visually compelling local content, and editable instance-level assets while preserving a consistent global terrain structure.