정지 상태 관찰로부터의 관절 객체 재구성
Articulated Object Reconstruction from Rest-State Observation
July 30, 2026
저자: Daeun Lee, Jaeah Lee, Woosung Kim, Haebeom Jung, Jaesik Park
cs.AI
초록
대화형 디지털 트윈을 구축하려면 객체가 어떻게 관절 운동을 수행하는지를 결정하는 3D 기하 구조와 운동학적 구조를 모두 복원해야 한다. 그러나 기존의 관절 객체 재구성 방법은 여러 관절 상태에서 명시적으로 관찰 가능한 움직임을 요구한다. 우리는 단일 닫힌 상태에서 관절 객체를 재구성하는 정지 상태 기반 정식화를 도입한다. 이는 기하·의미·움직임 사전이 움직임 단서의 부재를 보완해야 하는 본질적으로 ill-posed한 설정이다. 본 프레임워크는 명시적 메시를 교차 모델 검증 및 융합을 위한 중간 표현으로 사용하여, 비전-언어 모델과 분할 모델의 잡음이 많은 출력을 공간적으로 일관된 부품 구조로 조정한다. 관찰된 움직임 없이 관절 파라미터를 추정하기 위해, 비디오 확산 모델을 활용하여 관절 운동 가설을 합성하고 기하학적 일관성을 통해 이를 검증한다. 우리의 접근 방식은 정확한 부품 분해와 물리적으로 타당한 관절 운동을 달성하며, 움직임 관찰 기반 재구성 방법, 생성 기반 방법, 모듈형 사전 훈련 모델 기준선과 경쟁력 있는 성능을 보인다.
English
Building interactive digital twins requires recovering both 3D geometry and the kinematic structures that govern how objects articulate. Yet existing methods for articulated object reconstruction require explicitly observable motion from multiple articulation states. We introduce a rest-state formulation that reconstructs articulated objects from a single closed configuration, an inherently ill-posed setting where geometry, semantics, and motion priors compensate for the absence of motion cues. Our framework adopts an explicit mesh as an intermediate representation for cross-model verification and fusion, reconciling noisy outputs from vision-language and segmentation models into spatially consistent part structures. To estimate joint parameters without observed motion, we use a video diffusion model to synthesize articulation hypotheses and validate them through geometric consistency. Our approach achieves accurate part decomposition and physically plausible articulation, performing competitively with motion-observing reconstruction-based, generation-based, and modular pretrained-model baselines.