ChatPaper.aiChatPaper

HarmoHOI:协调外观与三维运动的多视角手物交互合成

HarmoHOI: Harmonizing Appearance and 3D Motion for Multi-view Hand-Object Interaction Synthesis

July 19, 2026
作者: Lingwei Dang, Juntong Li, Zonghan Li, Hongwen Zhang, Liang An, Wei Min, Yebin Liu, Qingyao Wu
cs.AI

摘要

手-物交互(HOI)合成是动画制作和具身智能的基石。尽管视频基础模型具备强大的先验知识,但由于手部运动复杂且存在遮挡,多视角一致的HOI合成仍具挑战性。我们提出HarmoHOI,一个统一的扩散框架,能够联合且和谐地生成同步的多视角HOI视频以及全局对齐的3D点轨迹。核心思路在于:鲁棒的多视角一致性本质上需要全局对齐的3D几何与运动。为此,我们设计了多视角扩散Transformer混合模型,该模型同时对RGB视频和3D点轨迹进行联合建模。通过将点轨迹表示为伪视频,我们将3D几何信号与基础模型的2D潜在空间对齐,从而最小化域差距并简化先验知识的适配。为进一步保证几何一致性,我们引入全局运动对齐扩散机制,将粗糙的点轨迹细化为具有度量尺度且全局对齐的3D轨迹。HarmoHOI在去噪过程中实现了2D外观与3D运动的即时协同演化。针对多视角HOI数据稀缺的问题,我们采用混合数据课程学习策略,成功将单视角数据中的通用先验迁移至同步多视角生成任务。实验结果表明,HarmoHOI在视觉质量、运动合理性和多视角几何一致性方面均达到当前最优水平。项目页面访问:https://droliven.github.io/HarmoHOI_project。
English
Hand-Object Interaction (HOI) synthesis is a cornerstone for animation production and embodied AI. Despite the strong priors of video foundation models, multi-view consistent HOI synthesis remains challenging due to complex hand motions and occlusions. We present HarmoHOI, a unified diffusion framework that jointly and harmoniously generates synchronized multi-view HOI videos and globally aligned 3D point tracks. Our core insight is that robust multi-view consistency fundamentally requires globally aligned 3D geometry and motion. To this end, we propose a Mixture of Multi-view Diffusion Transformer that co-models RGB videos and 3D point tracks. By representing point tracks as pseudo-videos, we align 3D geometric signals with the 2D latent space of foundation models, thereby minimizing the domain gap and easing adaptation of priors. To further ensure geometry consistency, we introduce Global Motion Aligning Diffusion, which refines coarse point tracks into metric-scale, globally aligned 3D trajectories. HarmoHOI enables on-the-fly co-evolution of 2D appearance and 3D motion during denoising. To overcome the scarcity of multi-view HOI data, we employ a hybrid data curriculum learning strategy that successfully transfers generic priors from single-view data to synchronized multi-view generation. Experimental results show that HarmoHOI achieves state-of-the-art performance in visual quality, motion plausibility, and multi-view geometric consistency. Project page available at https://droliven.github.io/HarmoHOI_project.