BLARM:透過混合潛在剛體運動基本單元,從影片製作3D物體動畫
BLARM: Animating 3D Objects from Video via Blending Latent Rigid Motion Primitives
August 31, 2026
作者: Pradyumn Goyal, Yizhak Ben-Shabat, Hsueh-Ti Derek Liu, Haomiao Jiang, Snehasish Mukherjee, Kyle Spence, Mark Stauber, Evangelos Kalogerakis, Yunze Zeng
cs.AI
摘要
我們提出BLARM,一種用於影片驅動的3D網格動畫的前饋方法。給定單目影片和靜態物件網格,BLARM預測一個時間上連貫的動畫網格,其動作跟隨影片。我們不依賴顯式綁定或直接回歸高維頂點運動,而是使用一組緊湊的、學習到的隨時間變化的剛體運動分量,以及不隨時間變化的頂點到分量蒙皮權重來表示動畫。這產生了一個低維度變形空間,無需骨架、籠、蒙皮權重或綁定標註。我們的架構透過分解的空間-時間注意力,以影片特徵對幾何推導出的變形潛在特徵進行條件化,然後解碼由預測蒙皮權重混合的剛體變換。透過軌跡重建、熵正則化和動作感知對比學習進行訓練,BLARM在恢復緊湊、可解釋的動作結構的同時,從單目影片中產生準確且時間穩定的動畫。
English
We introduce BLARM, a feed-forward method for video-driven 3D mesh animation. Given a monocular video and a static object mesh, BLARM predicts a temporally coherent animated mesh whose motion follows the video. Rather than relying on explicit rigs or directly regressing high-dimensional vertex motion, we represent animation using a compact set of learned, time-varying rigid motion components and time-invariant vertex-to-component skinning weights. This yields a low-dimensional deformation space without requiring skeletons, cages, skinning weights, or rig annotations. Our architecture conditions geometry-derived deformation latents on video features through factorized spatial-temporal attention, then decodes rigid transformations blended by predicted skinning weights. Trained with trajectory reconstruction, entropy regularization, and motion-aware contrastive learning, BLARM produces accurate and temporally stable animations while recovering compact, interpretable motion structure from monocular video.