HarmoHOI: 多視点手物体インタラクション合成のための外観と3D動作の調和
HarmoHOI: Harmonizing Appearance and 3D Motion for Multi-view Hand-Object Interaction Synthesis
July 19, 2026
著者: Lingwei Dang, Juntong Li, Zonghan Li, Hongwen Zhang, Liang An, Wei Min, Yebin Liu, Qingyao Wu
cs.AI
要旨
手と物体のインタラクション(HOI)合成は、アニメーション制作や具現化AIの基盤技術である。映像基盤モデルの強力な事前知識にもかかわらず、複雑な手の動きや遮蔽により、多視点一貫性のあるHOI合成は依然として困難である。我々は、同期した多視点HOI映像とグローバルに整合された3D点軌跡を共同かつ調和的に生成する統一拡散フレームワークHarmoHOIを提案する。我々の核心的な洞察は、ロバストな多視点一貫性の実現には、基本的にグローバルに整合された3D幾何と動作が必要であるという点である。このために、RGB映像と3D点軌跡を共にモデル化するMixture of Multi-view Diffusion Transformerを提案する。点軌跡を疑似映像として表現することで、3D幾何信号を基盤モデルの2D潜在空間に整合させ、ドメインギャップを最小化し、事前知識の適応を容易にする。さらに幾何的一貫性を確保するために、粗い点軌跡をメートルスケールでグローバルに整合された3D軌跡に洗練するGlobal Motion Aligning Diffusionを導入する。HarmoHOIは、ノイズ除去中に2D外観と3D動作の即時共進化を可能にする。多視点HOIデータの不足を克服するため、ハイブリッドデータカリキュラム学習戦略を採用し、単視点データからの汎用的事前知識を同期多視点生成にうまく転送する。実験結果は、HarmoHOIが視覚品質、動作の妥当性、多視点幾何一貫性において最先端の性能を達成することを示している。プロジェクトページはhttps://droliven.github.io/HarmoHOI_projectで公開されている。
English
Hand-Object Interaction (HOI) synthesis is a cornerstone for animation production and embodied AI. Despite the strong priors of video foundation models, multi-view consistent HOI synthesis remains challenging due to complex hand motions and occlusions. We present HarmoHOI, a unified diffusion framework that jointly and harmoniously generates synchronized multi-view HOI videos and globally aligned 3D point tracks. Our core insight is that robust multi-view consistency fundamentally requires globally aligned 3D geometry and motion. To this end, we propose a Mixture of Multi-view Diffusion Transformer that co-models RGB videos and 3D point tracks. By representing point tracks as pseudo-videos, we align 3D geometric signals with the 2D latent space of foundation models, thereby minimizing the domain gap and easing adaptation of priors. To further ensure geometry consistency, we introduce Global Motion Aligning Diffusion, which refines coarse point tracks into metric-scale, globally aligned 3D trajectories. HarmoHOI enables on-the-fly co-evolution of 2D appearance and 3D motion during denoising. To overcome the scarcity of multi-view HOI data, we employ a hybrid data curriculum learning strategy that successfully transfers generic priors from single-view data to synchronized multi-view generation. Experimental results show that HarmoHOI achieves state-of-the-art performance in visual quality, motion plausibility, and multi-view geometric consistency. Project page available at https://droliven.github.io/HarmoHOI_project.