BLARM: 潜在的剛体運動プリミティブのブレンディングによるビデオからの3Dオブジェクトのアニメーション
BLARM: Animating 3D Objects from Video via Blending Latent Rigid Motion Primitives
August 31, 2026
著者: Pradyumn Goyal, Yizhak Ben-Shabat, Hsueh-Ti Derek Liu, Haomiao Jiang, Snehasish Mukherjee, Kyle Spence, Mark Stauber, Evangelos Kalogerakis, Yunze Zeng
cs.AI
要旨
我々は、ビデオ駆動の3Dメッシュアニメーションのためのフィードフォワード手法であるBLARMを提案する。単眼ビデオと静的オブジェクトメッシュが与えられると、BLARMは、ビデオの動きに追従する時間的に一貫したアニメーションメッシュを予測する。明示的なリグに依存したり、高次元の頂点モーションを直接回帰するのではなく、学習された時間変化する剛体運動成分のコンパクトな集合と、時間不変の頂点-成分間スキニング重みを用いてアニメーションを表現する。これにより、スケルトン、ケージ、スキニング重み、リグ注釈を必要とせずに、低次元の変形空間が得られる。我々のアーキテクチャは、分解された時空間アテンションを介して幾何学由来の変形潜在変数をビデオ特徴に条件付けし、次に予測されたスキニング重みでブレンドされた剛体変換をデコードする。軌道再構成、エントロピー正則化、およびモーションを考慮した対照学習を用いて訓練されたBLARMは、単眼ビデオからコンパクトで解釈可能なモーション構造を復元しつつ、正確で時間的に安定したアニメーションを生成する。
English
We introduce BLARM, a feed-forward method for video-driven 3D mesh animation. Given a monocular video and a static object mesh, BLARM predicts a temporally coherent animated mesh whose motion follows the video. Rather than relying on explicit rigs or directly regressing high-dimensional vertex motion, we represent animation using a compact set of learned, time-varying rigid motion components and time-invariant vertex-to-component skinning weights. This yields a low-dimensional deformation space without requiring skeletons, cages, skinning weights, or rig annotations. Our architecture conditions geometry-derived deformation latents on video features through factorized spatial-temporal attention, then decodes rigid transformations blended by predicted skinning weights. Trained with trajectory reconstruction, entropy regularization, and motion-aware contrastive learning, BLARM produces accurate and temporally stable animations while recovering compact, interpretable motion structure from monocular video.