ChatPaper.aiChatPaper

WorldSculpt: 接地されたビデオからの構成的世界の生成

WorldSculpt: Generating Compositional Worlds from Grounded Videos

September 4, 2026
著者: Muyao Niu, Jixuan He, Ruihan Yu, Lian Fu, Yonghao Yu, Zheng-Hui Huang, Yifan Zhan, Fengbo Lan, Yongtao Ge, Yinqiang Zheng, Kaipeng Zhang, Zhixiang Wang
cs.AI

要旨

我々は、数百個の物体を含む雑然としたシーンから、構成論的3D表現を生成する問題を研究する。目標は、ゲーム、AR/VR、シミュレーション、ロボティクスといった下流アプリケーションで要求されるように、シーンを共有ワールド座標系内に配置された個々の物体メッシュの集合として表現することである。このタスクは、物体同士が強く遮蔽し合い、各視点からはその形状の一部しか観測できない密集したシーンにおいて困難を伴う。幾何学ベースの手法は通常、シーンを単一の表現として再構成し、遮蔽領域には不完全な形状を残す。一方、生成事前分布を利用する既存の構成論的手法は、比較的単純なシーンに大きく限られている。我々は、強力な単一物体3D生成事前分布を複数視点の観測に適応させることで、数百個の物体を含む複雑なシーンを代わりに構成論的に生成できることを示す。我々はこのパラダイムをPixal3Dによって具体化し、複数のポーズ付き観測に物体生成を基づかせる多視点条件付け経路を追加して拡張する。モデルは正準空間内の単一物体のみで完全にファインチューニングされているにもかかわらず、シーンレベルの学習を一切行わずに、深刻な遮蔽を含む大規模シーンへ一般化する。これは、このパラダイムの実現可能性とスケーラビリティを示すものである。さらに我々は、数百個の物体を含む密集した雑然シーン、物体ごとのアノテーション、およびグラウンドトゥルースメッシュからなるフォトリアリスティックなベンチマークであるUE-MeshySceneを導入する。単一物体、制御された複数物体、UE-MeshySceneの各評価において、本手法は既存手法を一貫して上回り、シーンの複雑さと遮蔽が増すにつれてその優位性はさらに大きくなる。最後に、MarbleやHY-World 2.0などの生成3DGSワールドを構成論的メッシュシーンへ変換することで、より広範な応用可能性を示す。
English
We study the problem of generating a compositional 3D representation of a cluttered scene containing hundreds of objects. The goal is to represent the scene as a collection of individual object meshes placed in a shared world frame, as required by downstream applications such as gaming, AR/VR, simulation, and robotics. This task is challenging in densely cluttered scenes, where objects heavily occlude one another and each view reveals only a fraction of their geometry. Geometry-based approaches typically reconstruct the scene as a single representation and leave incomplete geometry in occluded regions, while existing compositional methods with generative priors are largely limited to relatively simple scenes. We show that complex scenes with hundreds of objects can instead be generated compositionally by adapting a strong single-object 3D generative prior to multi-view observations. We instantiate this paradigm with Pixal3D, extending it with a multi-view conditioning pathway that grounds object generation in multiple posed observations. Although the model is finetuned entirely on single objects in canonical space, it generalizes to large scenes with severe occlusion without any scene-level training, demonstrating the feasibility and scalability of this paradigm. We further introduce UE-MeshyScene, a photorealistic benchmark of densely cluttered scenes with hundreds of objects, per-object annotations, and ground-truth meshes. Across single-object, controlled multi-object, and UE-MeshyScene evaluations, our method consistently outperforms prior approaches, with larger gains as scene complexity and occlusion increase. Finally, we demonstrate broader applicability by converting generated 3DGS worlds, such as Marble and HY-World 2.0, into compositional mesh scenes.