ChatPaper.aiChatPaper

HOMIE:基于多模态智能增强的以人和物体为中心的视频个性化

HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enchancement

July 20, 2026
作者: Yiyang Cai, Nan Chen, Rongchang Xie, Junwen Pan, Chunyang Jiang, Cheng Chen, Wen Zhou, Zhenbang Sun, Wei Xue, Wenhan Luo, Yike Guo
cs.AI

摘要

以人与物为中心的视频个性化(HOCVP)是主体驱动视频生成中的核心任务。然而,现有方法存在两个关键局限。首先,大多数专注于跨主体个性化的方法仍难以在主体保真度与人类与多样化物体(尤其是当物体代表抽象概念如标志时)之间的精确交互模式之间取得平衡。其次,尽管主体内参考(例如OCR映射、多视角输入)有望增强主体保真度,但大多数现有作品缺乏理解此类潜在对应关系的机制。为解决这两个挑战,我们提出HOMIE,一个以统一方式处理跨主体和主体内输入设置的HOCVP框架。与先前方法相比,HOMIE提出了一种更优的多模态大语言模型(MLLM)集成策略,以在不损害文本编码器可控性或不引发昂贵重新对齐的情况下,提取参考级关系的知识。具体而言,我们在自注意力中引入全局多模态引导,以更好地将MLLM导出的语义特征与VAE令牌对齐。此外,我们提出模态参考嵌入,以区分来自MLLM特征和VAE令牌的令牌,并关联主体内参考图像令牌。大量实验验证了我们的方法在各项HOCVP任务中均达到了最先进性能。项目页面:https://yiyangcai.github.io/homie-page.github.io/
English
Human-object centric video personalization (HOCVP) is a core task within subject-driven video generation. However, existing methods suffer from two key limitations. First, most approaches focusing on inter-subject personalization still struggle to strike a balance between high subject fidelity and accurate interaction patterns between humans and diverse objects, especially when objects represent abstract concepts such as logos. Second, while intra-subject references (e.g., OCR maps, multi-view inputs) are expected to enhance subject fidelity, most existing works lack mechanisms to understand such latent correspondence. To address both challenges, we propose HOMIE, an HOCVP framework that tackles both inter- and intra-subject input settings in a unified manner. Compared to previous approaches, HOMIE proposes a better MLLM integration strategy to extract knowledge of reference-level relationships without compromising the controllability of text encoders or incurring costly re-alignment. Specifically, we introduce global multimodal guidance within self-attention to better align MLLM-derived semantic features with VAE tokens. Furthermore, we propose modality-reference embedding to differentiate tokens from MLLM features and VAE tokens and associate intra-subject reference image tokens. Extensive experiments validate that our method achieves state-of-the-art performance across various HOCVP tasks. Project Page: https://yiyangcai.github.io/homie-page.github.io/