基于静止状态观测的铰接物体重建
Articulated Object Reconstruction from Rest-State Observation
July 30, 2026
作者: Daeun Lee, Jaeah Lee, Woosung Kim, Haebeom Jung, Jaesik Park
cs.AI
摘要
构建交互式数字孪生需要同时恢复三维几何结构以及控制物体关节运动的运动学结构。然而,现有的铰接物体重建方法要求从多个关节状态中显式观测运动信息。我们提出了一种静息状态公式化方法,能够从单一闭合构型中重建铰接物体——这是一个本质上不适定的设定,需要依赖几何、语义和运动先验来弥补运动线索的缺失。我们的框架采用显式网格作为中间表示,以实现跨模型验证与融合,将视觉-语言模型和分割模型产生的噪声输出协调为空间一致的部件结构。为了在无观测运动的情况下估计关节参数,我们使用视频扩散模型合成关节运动假设,并通过几何一致性对其进行验证。我们的方法实现了精确的部件分解和物理上合理的关节运动,与基于运动观测的重建方法、生成方法以及模块化预训练模型基线相比,性能具有竞争力。
English
Building interactive digital twins requires recovering both 3D geometry and the kinematic structures that govern how objects articulate. Yet existing methods for articulated object reconstruction require explicitly observable motion from multiple articulation states. We introduce a rest-state formulation that reconstructs articulated objects from a single closed configuration, an inherently ill-posed setting where geometry, semantics, and motion priors compensate for the absence of motion cues. Our framework adopts an explicit mesh as an intermediate representation for cross-model verification and fusion, reconciling noisy outputs from vision-language and segmentation models into spatially consistent part structures. To estimate joint parameters without observed motion, we use a video diffusion model to synthesize articulation hypotheses and validate them through geometric consistency. Our approach achieves accurate part decomposition and physically plausible articulation, performing competitively with motion-observing reconstruction-based, generation-based, and modular pretrained-model baselines.