WithEveryone: グループ画像生成のための統合プランニングとアイデンティティ接地
WithEveryone: Unified Planning and Identity Grounding for Group Image Generation
August 20, 2026
著者: Hengyuan Xu, Qixun Wang, Yiji Cheng, Miles Yang, Zhao Zhong, Wei Cheng, Xingjun Ma, Yu-gang Jiang
cs.AI
要旨
多数の指定人物をシーンに含めなければならない場合、アイデンティティ保持画像生成はますます信頼性が低下する。各アイデンティティを保持するだけでなく、モデルはすべての参照を異なる人物と位置に結び付ける必要があり、学習時のアイデンティティ損失は、ノイズを含む複数の予測顔の間で対応関係を確立しなければならない。我々は、最大10個の参照アイデンティティを含むグループ画像を生成するための統一フレームワークであるWithEveryoneを提案する。WithEveryoneは、選択された各アイデンティティをアドレス指定されたトークンとして注入し、構造化されたアイデンティティ-レイアウト計画を予測し、その計画を視覚的条件としてレンダリングする。その主要な目的関数であるLayout-Grounded ID Lossは、注釈付き顔領域を用いて意図したアイデンティティを直接教師信号として利用し、不安定な埋め込みベースの顔マッチングを回避する。ID Representation Forcingはさらに、画像合成の前に各アイデンティティに対する予測を学習させる。アイデンティティ非重複ベンチマークにおいて、WithEveryoneは最高のターゲットコンテキストにおけるアイデンティティ類似度を達成し、顔類似度をGPT-Image-2の0.462から0.499に改善するとともに、コピー&ペーストアーティファクトを0.169から0.055に低減する。さらに、要求されたアイデンティティの97.3%をカバーし、重複率はわずか2.8%である。これらの結果は、明示的なアイデンティティ-レイアウトのグラウンディングにより、直接の参照顔コピーに依存せずに、アイデンティティ保持生成をより大きなグループへ拡張できることを示している。
English
Identity-preserving image generation becomes increasingly unreliable when a scene must contain many specified people. Beyond retaining each identity, the model must bind every reference to a distinct person and location, while training-time identity losses must establish correspondence among several noisy predicted faces. We introduce WithEveryone, a unified framework for generating group images up to ten reference identities. WithEveryone injects each selected identity as an addressed token, predicts a structured identity--layout plan, and renders the plan as a visual condition. Its key objective, Layout-Grounded ID Loss, uses annotated face regions to supervise the intended identities directly, avoiding unstable embedding-based face matching; ID Representation Forcing additionally trains a prediction for each identity before image synthesis. On an identity-disjoint benchmark, WithEveryone achieves the highest target-context identity similarity, improving face similarity from 0.462 for GPT-Image-2 to 0.499, while reducing copy-paste artifacts from 0.169 to 0.055. It further covers 97.3\% of the requested identities with a duplicate rate of only 2.8\%. These results show that explicit identity--layout grounding enables identity-preserving generation to scale to larger groups without relying on direct reference-face copying.