ChatPaper.aiChatPaper

HarmoHOI: 다중 뷰 손-객체 상호작용 합성을 위한 외형과 3D 모션의 조화

HarmoHOI: Harmonizing Appearance and 3D Motion for Multi-view Hand-Object Interaction Synthesis

July 19, 2026
저자: Lingwei Dang, Juntong Li, Zonghan Li, Hongwen Zhang, Liang An, Wei Min, Yebin Liu, Qingyao Wu
cs.AI

초록

손-객체 상호작용(HOI) 합성은 애니메이션 제작과 구현형 AI의 핵심 요소이다. 비디오 기초 모델의 강력한 사전 조건에도 불구하고, 복잡한 손 움직임과 가려짐으로 인해 다중 시점 일관된 HOI 합성은 여전히 어려운 과제로 남아 있다. 본 논문에서는 동기화된 다중 시점 HOI 비디오와 전역적으로 정렬된 3D 포인트 트랙을 공동으로 조화롭게 생성하는 통합 확산 프레임워크인 HarmoHOI를 제안한다. 핵심 통찰은 강건한 다중 시점 일관성을 위해서는 근본적으로 전역적으로 정렬된 3D 형상과 움직임이 필요하다는 점이다. 이를 위해 RGB 비디오와 3D 포인트 트랙을 공동 모델링하는 다중 시점 확산 트랜스포머 혼합(Mixture of Multi-view Diffusion Transformer)을 제안한다. 포인트 트랙을 의사 비디오로 표현함으로써 3D 기하 신호를 기초 모델의 2D 잠재 공간과 정렬시켜 도메인 차이를 최소화하고 사전 조건의 적응을 용이하게 한다. 형상 일관성을 더욱 보장하기 위해 전역 움직임 정렬 확산(Global Motion Aligning Diffusion)을 도입하여 대략적인 포인트 트랙을 미터 단위의 전역 정렬된 3D 궤적으로 정제한다. HarmoHOI는 잡음 제거 과정에서 2D 외관과 3D 움직임의 실시간 공동 진화를 가능하게 한다. 다중 시점 HOI 데이터의 부족 문제를 극복하기 위해 하이브리드 데이터 커리큘럼 학습 전략을 채택하여 단일 시점 데이터의 일반적 사전 조건을 동기화된 다중 시점 생성으로 성공적으로 전이한다. 실험 결과, HarmoHOI는 시각적 품질, 움직임 타당성 및 다중 시점 기하 일관성 측면에서 최첨단 성능을 달성함을 보여준다. 프로젝트 페이지는 https://droliven.github.io/HarmoHOI_project 에서 확인할 수 있다.
English
Hand-Object Interaction (HOI) synthesis is a cornerstone for animation production and embodied AI. Despite the strong priors of video foundation models, multi-view consistent HOI synthesis remains challenging due to complex hand motions and occlusions. We present HarmoHOI, a unified diffusion framework that jointly and harmoniously generates synchronized multi-view HOI videos and globally aligned 3D point tracks. Our core insight is that robust multi-view consistency fundamentally requires globally aligned 3D geometry and motion. To this end, we propose a Mixture of Multi-view Diffusion Transformer that co-models RGB videos and 3D point tracks. By representing point tracks as pseudo-videos, we align 3D geometric signals with the 2D latent space of foundation models, thereby minimizing the domain gap and easing adaptation of priors. To further ensure geometry consistency, we introduce Global Motion Aligning Diffusion, which refines coarse point tracks into metric-scale, globally aligned 3D trajectories. HarmoHOI enables on-the-fly co-evolution of 2D appearance and 3D motion during denoising. To overcome the scarcity of multi-view HOI data, we employ a hybrid data curriculum learning strategy that successfully transfers generic priors from single-view data to synchronized multi-view generation. Experimental results show that HarmoHOI achieves state-of-the-art performance in visual quality, motion plausibility, and multi-view geometric consistency. Project page available at https://droliven.github.io/HarmoHOI_project.