ChatPaper.aiChatPaper

WorldClaw:大规模智能体式3D开放世界生成

WorldClaw: Agentic 3D Open-World Generation at Scale

August 5, 2026
作者: Chunchao Guo, Jinpeng Li, Yang Li, Zilong Huang
cs.AI

摘要

从开放式文本生成大规模、可自由探索的3D世界仍然是一项挑战,因为系统必须同时保持全局空间连贯性、丰富的局部内容,以及适合下游编辑和复用的显式资源。我们提出了WorldClaw,一个完全智能体驱动、由粗到细的开放世界3D场景生成框架。规划智能体将文本提示转换为关于区域、地形、资源、材质和空间关系的结构化规范。随后,WorldClaw通过语义布局、可复用资源、生成式或程序化材质以及区域感知高度场构建全局连贯的地形基础。对于细节要求高的区域,它生成地形条件化的组合,重建可编辑的带纹理网格,并恢复其在地形上的放置;基于渲染的智能体进一步细化地形、物体、外观和接触关系。在多样化的开放世界提示下,WorldClaw生成的大规模场景具有连贯的空间组织、视觉上引人入胜的局部内容,以及可编辑的实例级资源,同时保持一致的全局地形结构。
English
Generating large-scale, freely explorable 3D worlds from open-ended text remains challenging because a system must jointly maintain global spatial coherence, rich local content, and explicit assets suitable for downstream editing and reuse. We present WorldClaw, a fully agentic, coarse-to-fine framework for open-world 3D scene generation. Planning agents translate a text prompt into a structured specification of regions, terrain, assets, materials, and spatial relations. WorldClaw then builds a globally coherent terrain foundation from semantic layouts, reusable assets, generative or procedural materials, and a region-aware height field. For detail-demanding regions, it generates terrain-conditioned compositions, reconstructs editable textured meshes, and recovers their placement on the terrain; render-based agents further refine terrain, objects, appearance, and contacts. Across diverse open-world prompts, WorldClaw produces large-scale scenes with coherent spatial organization, visually compelling local content, and editable instance-level assets while preserving a consistent global terrain structure.