ChatPaper.aiChatPaper

WorldSculpt:從具象影片生成組合式世界

WorldSculpt: Generating Compositional Worlds from Grounded Videos

September 4, 2026
作者: Muyao Niu, Jixuan He, Ruihan Yu, Lian Fu, Yonghao Yu, Zheng-Hui Huang, Yifan Zhan, Fengbo Lan, Yongtao Ge, Yinqiang Zheng, Kaipeng Zhang, Zhixiang Wang
cs.AI

摘要

我們研究生成包含數百個物體的雜亂場景之組合式3D表示問題。目標是將場景表示為置於共享世界座標系中的個別物體網格集合,此為遊戲、AR/VR、模擬與機器人等下游應用所需之形式。此任務在密集雜亂場景中尤具挑戰性,物體彼此嚴重遮擋,每個視角僅能呈現其幾何形狀的一小部分。基於幾何的方法通常將場景重建為單一表示,並在遮擋區域留下不完整的幾何結構;而現有結合生成先驗的組合式方法大多僅適用於相對簡單的場景。我們表明,包含數百個物體的複雜場景可藉由將強大的單物體3D生成先驗適應至多視圖觀測,以組合式方式生成。我們以Pixal3D實現此範式,並擴展其多視圖條件路徑,使物體生成奠基於多個具姿態的觀測。儘管該模型完全在規範空間中的單物體上進行微調,它在不需任何場景級訓練的情況下,即能泛化至具有嚴重遮擋的大型場景,證明此範式的可行性與可擴展性。我們進一步引入UE-MeshyScene,一個涵蓋數百個物體、逐物體標註與真實網格的密集雜亂場景之逼真基準。在單物體、受控多物體及UE-MeshyScene的評估中,我們的方法持續優於既有方法,且於場景複雜度與遮擋程度增加時獲得更大優勢。最後,我們將生成的3DGS世界(如Marble和HY-World 2.0)轉換為組合式網格場景,以展示更廣泛的應用性。
English
We study the problem of generating a compositional 3D representation of a cluttered scene containing hundreds of objects. The goal is to represent the scene as a collection of individual object meshes placed in a shared world frame, as required by downstream applications such as gaming, AR/VR, simulation, and robotics. This task is challenging in densely cluttered scenes, where objects heavily occlude one another and each view reveals only a fraction of their geometry. Geometry-based approaches typically reconstruct the scene as a single representation and leave incomplete geometry in occluded regions, while existing compositional methods with generative priors are largely limited to relatively simple scenes. We show that complex scenes with hundreds of objects can instead be generated compositionally by adapting a strong single-object 3D generative prior to multi-view observations. We instantiate this paradigm with Pixal3D, extending it with a multi-view conditioning pathway that grounds object generation in multiple posed observations. Although the model is finetuned entirely on single objects in canonical space, it generalizes to large scenes with severe occlusion without any scene-level training, demonstrating the feasibility and scalability of this paradigm. We further introduce UE-MeshyScene, a photorealistic benchmark of densely cluttered scenes with hundreds of objects, per-object annotations, and ground-truth meshes. Across single-object, controlled multi-object, and UE-MeshyScene evaluations, our method consistently outperforms prior approaches, with larger gains as scene complexity and occlusion increase. Finally, we demonstrate broader applicability by converting generated 3DGS worlds, such as Marble and HY-World 2.0, into compositional mesh scenes.