ChatPaper.aiChatPaper

HOMIE: 멀티모달 지능형 향상을 통한 인간-객체 중심 비디오 개인화

HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enchancement

July 20, 2026
저자: Yiyang Cai, Nan Chen, Rongchang Xie, Junwen Pan, Chunyang Jiang, Cheng Chen, Wen Zhou, Zhenbang Sun, Wei Xue, Wenhan Luo, Yike Guo
cs.AI

초록

인간-객체 중심 비디오 개인화(HOCVP)는 주체 기반 비디오 생성(subject-driven video generation)의 핵심 과제입니다. 그러나 기존 방법들은 두 가지 주요 한계를 가집니다. 첫째, 주체 간 개인화(inter-subject personalization)에 초점을 맞춘 대부분의 접근법은 높은 주체 충실도와 인간과 다양한 객체 간의 정확한 상호작용 패턴 사이에서 균형을 맞추는 데 여전히 어려움을 겪습니다. 특히 객체가 로고와 같은 추상적 개념을 나타낼 때 더욱 그렇습니다. 둘째, 주체 내 참조(intra-subject reference, 예: OCR 맵, 다중 뷰 입력)가 주체 충실도를 향상시킬 것으로 기대되지만, 대부분의 기존 연구는 이러한 잠재적 대응 관계를 이해하는 메커니즘이 부족합니다. 이 두 가지 문제를 해결하기 위해, 우리는 HOMIE를 제안합니다. 이는 주체 간 및 주체 내 입력 설정을 통합된 방식으로 처리하는 HOCVP 프레임워크입니다. 이전 접근법과 비교하여 HOMIE는 텍스트 인코더의 제어 가능성을 손상시키거나 값비싼 재정렬을 유발하지 않으면서 참조 수준의 관계에 대한 지식을 추출하는 더 나은 MLLM 통합 전략을 제안합니다. 구체적으로, 우리는 셀프 어텐션 내에서 전역 다중 모달 가이던스를 도입하여 MLLM에서 도출된 의미적 특징을 VAE 토큰과 더 잘 정렬합니다. 또한, 모달리티-참조 임베딩(modality-reference embedding)을 제안하여 MLLM 특징과 VAE 토큰의 토큰을 구별하고 주체 내 참조 이미지 토큰을 연관시킵니다. 광범위한 실험을 통해 우리의 방법이 다양한 HOCVP 과제에서 최첨단 성능을 달성함을 입증했습니다. 프로젝트 페이지: https://yiyangcai.github.io/homie-page.github.io/
English
Human-object centric video personalization (HOCVP) is a core task within subject-driven video generation. However, existing methods suffer from two key limitations. First, most approaches focusing on inter-subject personalization still struggle to strike a balance between high subject fidelity and accurate interaction patterns between humans and diverse objects, especially when objects represent abstract concepts such as logos. Second, while intra-subject references (e.g., OCR maps, multi-view inputs) are expected to enhance subject fidelity, most existing works lack mechanisms to understand such latent correspondence. To address both challenges, we propose HOMIE, an HOCVP framework that tackles both inter- and intra-subject input settings in a unified manner. Compared to previous approaches, HOMIE proposes a better MLLM integration strategy to extract knowledge of reference-level relationships without compromising the controllability of text encoders or incurring costly re-alignment. Specifically, we introduce global multimodal guidance within self-attention to better align MLLM-derived semantic features with VAE tokens. Furthermore, we propose modality-reference embedding to differentiate tokens from MLLM features and VAE tokens and associate intra-subject reference image tokens. Extensive experiments validate that our method achieves state-of-the-art performance across various HOCVP tasks. Project Page: https://yiyangcai.github.io/homie-page.github.io/