ChatPaper.aiChatPaper

テキスト画像個人化モデルにおける潜在アイデンティティチューニング

Latent-Identity Tuning in Text-to-Image Personalization Models

July 13, 2026
著者: Daniel Garibi, Ronen Kamenetsky, Hadar Averbuch-Elor, Daniel Cohen-Or, Or Patashnik
cs.AI

要旨

人の顔の生成や編集には高度な精度が求められる。わずかな変更でも、被写体の知覚される同一性が大きく変わることがあるからだ。しかし、汎用的なテキスト画像生成モデルを基にした現在の個人化手法や編集手法では、顔の細かい編集に必要な精度が不足していることが多い。本稿では、テキスト画像生成における個人化モデルに対して、細粒度の同一性チューニングを実現する手法を提案する。通常の画像編集が与えられた画像に対して処理を行うのとは異なり、同一性チューニングは特定の人物の潜在表現を修正することで、編集後の同一性を一貫して保ちつつ多様な画像を生成できるようにする。この細粒度の潜在的な同一性チューニングを実現するため、テキスト画像個人化向けに事前学習済みで凍結されたエンコーダの潜在空間を探査する。本手法は追加の学習を必要とせず、代わりに凍結されたエンコーダの既存のアーキテクチャを活用して潜在的な意味的方向性を発見する。この空間は、人物の同一性のさまざまな側面を捉える上で異なる役割を担う一連の潜在トークンから構成され、これらは多くの場合、特定の空間的または意味的な顔領域に対応する。我々は、この空間内および選択したトークンによって定義される部分空間内で意味のある方向性を特定できることを示し、局所的で細粒度かつ意味的に一貫した編集を可能にする。本手法は、質的・量的実験により、画像間での同一性の一貫性を保ちながら多様な局所的な顔編集を実現することを検証する。プロジェクトページ: https://garibida.github.io/IdentityTuning/
English
Generating and editing a person's face demands high precision, as even minor modifications can significantly alter a subject's perceived identity. Current personalization and editing methods built on general-purpose text-to-image models, however, often lack the precision required for fine-grained facial edits. We present a method for fine-grained identity tuning in text-to-image personalization models. Unlike standard image editing, which operates on a given image, identity tuning modifies the latent representation of a specific identity, enabling the generation of diverse images that consistently depict the same edited identity. To enable fine-grained latent identity tuning, we explore the latent space of a pre-trained, frozen encoder for text-to-image personalization. Our approach requires no additional training. Instead, it leverages the existing architecture of a frozen encoder to uncover latent semantic directions. This space consists of a set of latent tokens that play distinct roles in capturing different aspects of an identity and often correspond to specific spatial or semantic facial regions. We show that meaningful directions can be identified within this space and within subspaces defined by selected tokens, enabling localized, fine-grained, and semantically coherent edits. We validate our approach through qualitative and quantitative experiments that demonstrate diverse localized facial edits while preserving cross-image identity consistency. Project page at: https://garibida.github.io/IdentityTuning/