ChatPaper.aiChatPaper

静止状態の観察に基づく関節物体の再構成

Articulated Object Reconstruction from Rest-State Observation

July 30, 2026
著者: Daeun Lee, Jaeah Lee, Woosung Kim, Haebeom Jung, Jaesik Park
cs.AI

要旨

インタラクティブなデジタルツインの構築には、3Dジオメトリと物体の関節運動を支配する運動学的構造の両方を復元することが必要である。しかし、既存の関節物体再構成手法は、複数の関節状態から明示的に観測可能な運動を必要とする。本研究では、単一の閉じた配置から関節物体を再構成する静止状態の定式化を導入する。これは、幾何学、意味論、および運動の事前知識が運動の手がかりの欠如を補完する、本質的に不良設定の状況である。本フレームワークは、クロスモデル検証と融合のための中間表現として明示的なメッシュを採用し、視覚言語モデルとセグメンテーションモデルからのノイズを含む出力を、空間的に整合性のあるパーツ構造へと統合する。観測された運動なしで関節パラメータを推定するため、ビデオ拡散モデルを用いて関節運動の仮説を合成し、幾何学的整合性を通じてそれらを検証する。本手法は、正確なパーツ分割と物理的に妥当な関節運動を実現し、運動を観測する再構成ベース、生成ベース、およびモジュール型事前学習モデルのベースラインと競争力のある性能を達成する。
English
Building interactive digital twins requires recovering both 3D geometry and the kinematic structures that govern how objects articulate. Yet existing methods for articulated object reconstruction require explicitly observable motion from multiple articulation states. We introduce a rest-state formulation that reconstructs articulated objects from a single closed configuration, an inherently ill-posed setting where geometry, semantics, and motion priors compensate for the absence of motion cues. Our framework adopts an explicit mesh as an intermediate representation for cross-model verification and fusion, reconciling noisy outputs from vision-language and segmentation models into spatially consistent part structures. To estimate joint parameters without observed motion, we use a video diffusion model to synthesize articulation hypotheses and validate them through geometric consistency. Our approach achieves accurate part decomposition and physically plausible articulation, performing competitively with motion-observing reconstruction-based, generation-based, and modular pretrained-model baselines.