WithEveryone:群體圖像生成的統一規劃與身份對齊
WithEveryone: Unified Planning and Identity Grounding for Group Image Generation
August 20, 2026
作者: Hengyuan Xu, Qixun Wang, Yiji Cheng, Miles Yang, Zhao Zhong, Wei Cheng, Xingjun Ma, Yu-gang Jiang
cs.AI
摘要
身份保持影像生成在場景必須包含多位指定人物時,其可靠性會大幅下降。除了保留每位人物的身份之外,模型還必須將每個參考影像綁定到特定人物及位置,同時訓練時的身份損失函數必須在多個雜訊預測人臉之間建立對應關係。我們提出 WithEveryone,一個統一框架,可生成最多包含十位參考身份的群體影像。WithEveryone 將每個選定身份注入為具位址的標記,預測結構化的身份–佈局計畫,並將該計畫渲染為視覺條件。其關鍵目標函式「佈局對齊身份損失」(Layout-Grounded ID Loss)利用標註的人臉區域直接監督預期的身份,避免了不穩定的基於嵌入的人臉匹配;「身份表徵強制」(ID Representation Forcing)則在影像合成之前進一步針對每個身份訓練預測。在身份不相交的基準測試上,WithEveryone 達到了最高的目標情境身份相似度,將臉部相似度從 GPT-Image-2 的 0.462 提升至 0.499,同時將複製貼上偽影從 0.169 降低至 0.055。此外,其對所請求身份的涵蓋率達 97.3%,重複率僅 2.8%。這些結果表明,明確的身份–佈局對齊能使身份保持生成擴展至更大的群體,而無需依賴直接參考人臉的複製。
English
Identity-preserving image generation becomes increasingly unreliable when a scene must contain many specified people. Beyond retaining each identity, the model must bind every reference to a distinct person and location, while training-time identity losses must establish correspondence among several noisy predicted faces. We introduce WithEveryone, a unified framework for generating group images up to ten reference identities. WithEveryone injects each selected identity as an addressed token, predicts a structured identity--layout plan, and renders the plan as a visual condition. Its key objective, Layout-Grounded ID Loss, uses annotated face regions to supervise the intended identities directly, avoiding unstable embedding-based face matching; ID Representation Forcing additionally trains a prediction for each identity before image synthesis. On an identity-disjoint benchmark, WithEveryone achieves the highest target-context identity similarity, improving face similarity from 0.462 for GPT-Image-2 to 0.499, while reducing copy-paste artifacts from 0.169 to 0.055. It further covers 97.3\% of the requested identities with a duplicate rate of only 2.8\%. These results show that explicit identity--layout grounding enables identity-preserving generation to scale to larger groups without relying on direct reference-face copying.