HOMIE: マルチモーダル知的強化による人間-物体中心のビデオパーソナライゼーション
HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enchancement
July 20, 2026
著者: Yiyang Cai, Nan Chen, Rongchang Xie, Junwen Pan, Chunyang Jiang, Cheng Chen, Wen Zhou, Zhenbang Sun, Wei Xue, Wenhan Luo, Yike Guo
cs.AI
要旨
人物-物体中心のビデオパーソナライゼーション(HOCVP)は、被写体駆動型ビデオ生成における中核的なタスクです。しかしながら、既存手法には2つの主要な制限があります。第一に、被写体間パーソナライゼーションに焦点を当てたほとんどの手法は、高い被写体忠実度と、人間と多様な物体(特にロゴなどの抽象的概念を表す物体)との間の正確な相互作用パターンのバランスを取るのに依然として苦労しています。第二に、被写体内参照(例えば、OCRマップ、多視点入力)は被写体忠実度を向上させることが期待されていますが、既存のほとんどの研究にはそのような潜在的な対応関係を理解するメカニズムが欠けています。これらの両方の課題に対処するため、我々はHOMIEを提案します。これは、被写体間および被写体内の両方の入力設定を統一的に扱うHOCVPフレームワークです。従来の手法と比較して、HOMIEは、テキストエンコーダの制御可能性を損なうことなく、またコストのかかる再調整を伴うことなく、参照レベルの関係性の知識を抽出するためのより優れたMLLM統合戦略を提案します。具体的には、自己注意内にグローバルなマルチモーダルガイダンスを導入し、MLLM由来の意味的特徴とVAEトークンをより良く整列させます。さらに、モダリティ参照埋め込みを提案し、MLLM特徴とVAEトークンからのトークンを区別し、被写体内参照画像トークンを関連付けます。広範な実験により、我々の手法が様々なHOCVPタスクにおいて最先端の性能を達成することが検証されました。Project Page: https://yiyangcai.github.io/homie-page.github.io/
English
Human-object centric video personalization (HOCVP) is a core task within subject-driven video generation. However, existing methods suffer from two key limitations. First, most approaches focusing on inter-subject personalization still struggle to strike a balance between high subject fidelity and accurate interaction patterns between humans and diverse objects, especially when objects represent abstract concepts such as logos. Second, while intra-subject references (e.g., OCR maps, multi-view inputs) are expected to enhance subject fidelity, most existing works lack mechanisms to understand such latent correspondence. To address both challenges, we propose HOMIE, an HOCVP framework that tackles both inter- and intra-subject input settings in a unified manner. Compared to previous approaches, HOMIE proposes a better MLLM integration strategy to extract knowledge of reference-level relationships without compromising the controllability of text encoders or incurring costly re-alignment. Specifically, we introduce global multimodal guidance within self-attention to better align MLLM-derived semantic features with VAE tokens. Furthermore, we propose modality-reference embedding to differentiate tokens from MLLM features and VAE tokens and associate intra-subject reference image tokens. Extensive experiments validate that our method achieves state-of-the-art performance across various HOCVP tasks. Project Page: https://yiyangcai.github.io/homie-page.github.io/