基於靜止狀態觀察的關節物體重建
Articulated Object Reconstruction from Rest-State Observation
July 30, 2026
作者: Daeun Lee, Jaeah Lee, Woosung Kim, Haebeom Jung, Jaesik Park
cs.AI
摘要
建構互動式數位孿生需要同時恢復物體的三維幾何與其關節運動所遵循的運動學結構。然而,現有的關節物體重建方法需要從多個關節狀態中明確觀察其運動。我們提出一種靜止態表述,可從單一閉合構型重建關節物體;此設定本質上是不適定的,需藉由幾何、語義與運動先驗來彌補運動線索的缺失。我們的框架採用顯式網格作為中間表示,以進行跨模型驗證與融合,將視覺-語言模型與分割模型所產生的雜訊輸出調和為空間一致的部件結構。為在無觀測運動之情況下估計關節參數,我們使用影片擴散模型合成關節運動假設,並透過幾何一致性加以驗證。我們的方法實現了精確的部件分解與物理上合理的關節運動,在效能上可與基於運動觀測的重建方法、生成方法及模組化預訓練模型基線並駕齊驅。
English
Building interactive digital twins requires recovering both 3D geometry and the kinematic structures that govern how objects articulate. Yet existing methods for articulated object reconstruction require explicitly observable motion from multiple articulation states. We introduce a rest-state formulation that reconstructs articulated objects from a single closed configuration, an inherently ill-posed setting where geometry, semantics, and motion priors compensate for the absence of motion cues. Our framework adopts an explicit mesh as an intermediate representation for cross-model verification and fusion, reconciling noisy outputs from vision-language and segmentation models into spatially consistent part structures. To estimate joint parameters without observed motion, we use a video diffusion model to synthesize articulation hypotheses and validate them through geometric consistency. Our approach achieves accurate part decomposition and physically plausible articulation, performing competitively with motion-observing reconstruction-based, generation-based, and modular pretrained-model baselines.