ChatPaper.aiChatPaper

BLARM: 비디오에서 잠재 강체 운동 프리미티브의 블렌딩을 통한 3D 객체 애니메이션

BLARM: Animating 3D Objects from Video via Blending Latent Rigid Motion Primitives

August 31, 2026
저자: Pradyumn Goyal, Yizhak Ben-Shabat, Hsueh-Ti Derek Liu, Haomiao Jiang, Snehasish Mukherjee, Kyle Spence, Mark Stauber, Evangelos Kalogerakis, Yunze Zeng
cs.AI

초록

우리는 비디오 기반 3D 메시 애니메이션을 위한 피드포워드 방법인 BLARM을 소개한다. 단안 비디오와 정적 객체 메시가 주어졌을 때, BLARM은 비디오의 움직임을 따르는 시간적으로 일관된 애니메이션 메시를 예측한다. 명시적 리그에 의존하거나 고차원 정점 운동을 직접 회귀하는 대신, 우리는 애니메이션을 학습된 시변 강체 운동 성분과 시불변 정점-성분 스키닝 가중치의 컴팩트한 집합으로 표현한다. 이는 골격, 케이지, 스키닝 가중치, 또는 리그 주석 없이 저차원 변형 공간을 생성한다. 우리의 아키텍처는 분해된 공간-시간 주의 메커니즘을 통해 기하학 기반 변형 잠재 변수를 비디오 특징에 조건화한 다음, 예측된 스키닝 가중치로 혼합된 강체 변환을 디코딩한다. 궤적 재구성, 엔트로피 정규화, 그리고 운동 인지 대조 학습으로 훈련된 BLARM은 단안 비디오로부터 컴팩트하고 해석 가능한 운동 구조를 복구하면서 정확하고 시간적으로 안정적인 애니메이션을 생성한다.
English
We introduce BLARM, a feed-forward method for video-driven 3D mesh animation. Given a monocular video and a static object mesh, BLARM predicts a temporally coherent animated mesh whose motion follows the video. Rather than relying on explicit rigs or directly regressing high-dimensional vertex motion, we represent animation using a compact set of learned, time-varying rigid motion components and time-invariant vertex-to-component skinning weights. This yields a low-dimensional deformation space without requiring skeletons, cages, skinning weights, or rig annotations. Our architecture conditions geometry-derived deformation latents on video features through factorized spatial-temporal attention, then decodes rigid transformations blended by predicted skinning weights. Trained with trajectory reconstruction, entropy regularization, and motion-aware contrastive learning, BLARM produces accurate and temporally stable animations while recovering compact, interpretable motion structure from monocular video.