BLARM: 通过混合潜在刚体运动基元从视频中驱动3D对象
BLARM: Animating 3D Objects from Video via Blending Latent Rigid Motion Primitives
August 31, 2026
作者: Pradyumn Goyal, Yizhak Ben-Shabat, Hsueh-Ti Derek Liu, Haomiao Jiang, Snehasish Mukherjee, Kyle Spence, Mark Stauber, Evangelos Kalogerakis, Yunze Zeng
cs.AI
摘要
我们提出BLARM,一种用于视频驱动3D网格动画的前馈方法。给定单目视频和静态物体网格,BLARM预测运动跟随视频的时间连贯动画网格。我们不依赖显式绑定或直接回归高维顶点运动,而是使用一组紧凑的学习得到的随时间变化的刚性运动分量和与时间无关的顶点到分量蒙皮权重来表示动画。这产生了低维形变空间,无需骨架、笼形控制器、蒙皮权重或绑定标注。我们的架构通过分解的时空注意力将几何导出的形变潜变量以视频特征为条件,然后解码由预测蒙皮权重混合的刚性变换。通过轨迹重建、熵正则化和运动感知对比学习进行训练,BLARM生成准确且时间稳定的动画,同时从单目视频中恢复紧凑、可解释的运动结构。
English
We introduce BLARM, a feed-forward method for video-driven 3D mesh animation. Given a monocular video and a static object mesh, BLARM predicts a temporally coherent animated mesh whose motion follows the video. Rather than relying on explicit rigs or directly regressing high-dimensional vertex motion, we represent animation using a compact set of learned, time-varying rigid motion components and time-invariant vertex-to-component skinning weights. This yields a low-dimensional deformation space without requiring skeletons, cages, skinning weights, or rig annotations. Our architecture conditions geometry-derived deformation latents on video features through factorized spatial-temporal attention, then decodes rigid transformations blended by predicted skinning weights. Trained with trajectory reconstruction, entropy regularization, and motion-aware contrastive learning, BLARM produces accurate and temporally stable animations while recovering compact, interpretable motion structure from monocular video.