ChatPaper.aiChatPaper

WorldClaw: 대규모 에이전트 기반 3D 오픈월드 생성

WorldClaw: Agentic 3D Open-World Generation at Scale

August 5, 2026
저자: Chunchao Guo, Jinpeng Li, Yang Li, Zilong Huang
cs.AI

초록

대규모의 자유 탐험 가능한 3D 세계를 개방형 텍스트로부터 생성하는 것은 시스템이 전역적 공간 일관성, 풍부한 로컬 콘텐츠, 그리고 후속 편집 및 재사용에 적합한 명시적 자산을 동시에 유지해야 하기 때문에 여전히 어려운 과제이다. 본 논문에서는 개방형 3D 장면 생성을 위한 완전한 에이전트 기반의 조대-정밀(coarse-to-fine) 프레임워크인 WorldClaw를 제시한다. 계획 에이전트는 텍스트 프롬프트를 영역, 지형, 자산, 머티리얼, 공간 관계에 대한 구조화된 명세로 변환한다. 이후 WorldClaw는 의미론적 레이아웃, 재사용 가능한 자산, 생성적 또는 절차적 머티리얼, 그리고 영역 인식 높이 필드로부터 전역적으로 일관된 지형 기반을 구축한다. 세부 묘사가 요구되는 영역의 경우, 지형 조건부 구성을 생성하고, 편집 가능한 텍스처 메시를 재구성하며, 지형 상에서의 배치를 복원한다. 렌더 기반 에이전트는 지형, 객체, 외관, 접촉을 추가로 정제한다. 다양한 개방형 세계 프롬프트에 걸쳐 WorldClaw는 일관된 전역 지형 구조를 유지하면서, 공간적으로 정합적인 구성을 갖춘 대규모 장면, 시각적으로 매력적인 로컬 콘텐츠, 그리고 편집 가능한 인스턴스 수준 자산을 생성한다.
English
Generating large-scale, freely explorable 3D worlds from open-ended text remains challenging because a system must jointly maintain global spatial coherence, rich local content, and explicit assets suitable for downstream editing and reuse. We present WorldClaw, a fully agentic, coarse-to-fine framework for open-world 3D scene generation. Planning agents translate a text prompt into a structured specification of regions, terrain, assets, materials, and spatial relations. WorldClaw then builds a globally coherent terrain foundation from semantic layouts, reusable assets, generative or procedural materials, and a region-aware height field. For detail-demanding regions, it generates terrain-conditioned compositions, reconstructs editable textured meshes, and recovers their placement on the terrain; render-based agents further refine terrain, objects, appearance, and contacts. Across diverse open-world prompts, WorldClaw produces large-scale scenes with coherent spatial organization, visually compelling local content, and editable instance-level assets while preserving a consistent global terrain structure.