ChatPaper.aiChatPaper

텍스트-이미지 개인화 모델에서의 잠재 정체성 튜닝

Latent-Identity Tuning in Text-to-Image Personalization Models

July 13, 2026
저자: Daniel Garibi, Ronen Kamenetsky, Hadar Averbuch-Elor, Daniel Cohen-Or, Or Patashnik
cs.AI

초록

사람의 얼굴 생성 및 편집은 높은 정밀도를 요구한다. 사소한 수정조차도 대상의 인식된 정체성을 크게 바꿀 수 있기 때문이다. 그러나 일반 목적의 텍스트-이미지 모델 기반의 현재 개인화 및 편집 방법은 세밀한 얼굴 편집에 필요한 정밀도가 부족한 경우가 많다. 본 논문에서는 텍스트-이미지 개인화 모델에서 세밀한 정체성 튜닝 방법을 제시한다. 주어진 이미지에 대해 작동하는 표준 이미지 편집과 달리, 정체성 튜닝은 특정 정체성의 잠재 표현을 수정하여 동일한 편집된 정체성을 일관되게 묘사하는 다양한 이미지 생성을 가능하게 한다. 세밀한 잠재 정체성 튜닝을 가능하게 하기 위해, 텍스트-이미지 개인화를 위해 사전 훈련되고 고정된 인코더의 잠재 공간을 탐색한다. 우리의 접근 방식은 추가 훈련이 필요하지 않다. 대신, 고정된 인코더의 기존 아키텍처를 활용하여 잠재 의미 방향을 발견한다. 이 공간은 정체성의 다양한 측면을 포착하는 데 뚜렷한 역할을 하며 종종 특정 공간적 또는 의미적 얼굴 영역에 해당하는 잠재 토큰 집합으로 구성된다. 우리는 이 공간 내에서 그리고 선택된 토큰에 의해 정의된 부분 공간 내에서 의미 있는 방향을 식별할 수 있음을 보여주며, 이를 통해 국소적이고 세밀하며 의미적으로 일관된 편집이 가능하다. 우리는 이미지 간 정체성 일관성을 유지하면서 다양한 국소적 얼굴 편집을 입증하는 정성적 및 정량적 실험을 통해 접근 방식을 검증한다. 프로젝트 페이지: https://garibida.github.io/IdentityTuning/
English
Generating and editing a person's face demands high precision, as even minor modifications can significantly alter a subject's perceived identity. Current personalization and editing methods built on general-purpose text-to-image models, however, often lack the precision required for fine-grained facial edits. We present a method for fine-grained identity tuning in text-to-image personalization models. Unlike standard image editing, which operates on a given image, identity tuning modifies the latent representation of a specific identity, enabling the generation of diverse images that consistently depict the same edited identity. To enable fine-grained latent identity tuning, we explore the latent space of a pre-trained, frozen encoder for text-to-image personalization. Our approach requires no additional training. Instead, it leverages the existing architecture of a frozen encoder to uncover latent semantic directions. This space consists of a set of latent tokens that play distinct roles in capturing different aspects of an identity and often correspond to specific spatial or semantic facial regions. We show that meaningful directions can be identified within this space and within subspaces defined by selected tokens, enabling localized, fine-grained, and semantically coherent edits. We validate our approach through qualitative and quantitative experiments that demonstrate diverse localized facial edits while preserving cross-image identity consistency. Project page at: https://garibida.github.io/IdentityTuning/