WorldClaw: 大規模なエージェント型3Dオープンワールド生成
WorldClaw: Agentic 3D Open-World Generation at Scale
August 5, 2026
著者: Chunchao Guo, Jinpeng Li, Yang Li, Zilong Huang
cs.AI
要旨
自由形式のテキストから大規模かつ自由に探索可能な3Dワールドを生成することは、システムがグローバルな空間的一貫性、豊かなローカルコンテンツ、および下流の編集・再利用に適した明示的なアセットを同時に維持する必要があるため、依然として困難な課題である。本稿では、オープンワールド3Dシーン生成のための完全エージェント型コース・ツー・ファインフレームワークであるWorldClawを提案する。プランニングエージェントは、テキストプロンプトを領域、地形、アセット、マテリアル、空間関係の構造化仕様に変換する。WorldClawは次に、セマンティックレイアウト、再利用可能なアセット、生成的または手続き型のマテリアル、および領域認識型高さ場から、グローバルに一貫した地形基盤を構築する。詳細を要求する領域に対しては、地形条件付き合成を生成し、編集可能なテクスチャ付きメッシュを再構成して、地形上での配置を復元する。さらに、レンダリングベースのエージェントが地形、オブジェクト、外観、および接触を精緻化する。多様なオープンワールドプロンプトにわたって、WorldClawは一貫したグローバル地形構造を維持しつつ、空間的に整合性のある構成、視覚的に魅力的なローカルコンテンツ、および編集可能なインスタンスレベルのアセットを備えた大規模シーンを生成する。
English
Generating large-scale, freely explorable 3D worlds from open-ended text remains challenging because a system must jointly maintain global spatial coherence, rich local content, and explicit assets suitable for downstream editing and reuse. We present WorldClaw, a fully agentic, coarse-to-fine framework for open-world 3D scene generation. Planning agents translate a text prompt into a structured specification of regions, terrain, assets, materials, and spatial relations. WorldClaw then builds a globally coherent terrain foundation from semantic layouts, reusable assets, generative or procedural materials, and a region-aware height field. For detail-demanding regions, it generates terrain-conditioned compositions, reconstructs editable textured meshes, and recovers their placement on the terrain; render-based agents further refine terrain, objects, appearance, and contacts. Across diverse open-world prompts, WorldClaw produces large-scale scenes with coherent spatial organization, visually compelling local content, and editable instance-level assets while preserving a consistent global terrain structure.