ChatPaper.aiChatPaper

HOMIE:基於多模態智能增強的人-物中心化視頻個性化

HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enchancement

July 20, 2026
作者: Yiyang Cai, Nan Chen, Rongchang Xie, Junwen Pan, Chunyang Jiang, Cheng Chen, Wen Zhou, Zhenbang Sun, Wei Xue, Wenhan Luo, Yike Guo
cs.AI

摘要

人類-物體中心視頻個性化(HOCVP)是主體驅動影片生成中的核心任務。然而,現有方法存在兩個主要侷限。首先,多數專注於主體間個性化的方法,仍難以在高主體保真度與人類與多樣物體間準確交互模式之間取得平衡,尤其是當物體代表抽象概念(如標誌)時。其次,雖然主體內參考(例如OCR圖、多視角輸入)預期能增強主體保真度,但現有大多數作品缺乏理解此類潛在對應關係的機制。為解決這兩項挑戰,我們提出HOMIE,一個統一處理主體間與主體內輸入設定的HOCVP框架。與先前方法相比,HOMIE提出更好的MLLM整合策略,以提取參考層級關係的知識,同時不損害文字編碼器的可控性,也無需承擔高昂的重新對齊成本。具體而言,我們在自注意力中引入全局多模態引導,以更佳地對齊MLLM導出的語義特徵與VAE標記。此外,我們提出模態參考嵌入,用以區分來自MLLM特徵與VAE標記的標記,並關聯主體內參考影像標記。廣泛的實驗驗證了我們的方法在各類HOCVP任務中達到最先進的效能。專案頁面:https://yiyangcai.github.io/homie-page.github.io/
English
Human-object centric video personalization (HOCVP) is a core task within subject-driven video generation. However, existing methods suffer from two key limitations. First, most approaches focusing on inter-subject personalization still struggle to strike a balance between high subject fidelity and accurate interaction patterns between humans and diverse objects, especially when objects represent abstract concepts such as logos. Second, while intra-subject references (e.g., OCR maps, multi-view inputs) are expected to enhance subject fidelity, most existing works lack mechanisms to understand such latent correspondence. To address both challenges, we propose HOMIE, an HOCVP framework that tackles both inter- and intra-subject input settings in a unified manner. Compared to previous approaches, HOMIE proposes a better MLLM integration strategy to extract knowledge of reference-level relationships without compromising the controllability of text encoders or incurring costly re-alignment. Specifically, we introduce global multimodal guidance within self-attention to better align MLLM-derived semantic features with VAE tokens. Furthermore, we propose modality-reference embedding to differentiate tokens from MLLM features and VAE tokens and associate intra-subject reference image tokens. Extensive experiments validate that our method achieves state-of-the-art performance across various HOCVP tasks. Project Page: https://yiyangcai.github.io/homie-page.github.io/