ChatPaper.aiChatPaper

WithEveryone:群体图像生成中的统一规划与身份锚定

WithEveryone: Unified Planning and Identity Grounding for Group Image Generation

August 20, 2026
作者: Hengyuan Xu, Qixun Wang, Yiji Cheng, Miles Yang, Zhao Zhong, Wei Cheng, Xingjun Ma, Yu-gang Jiang
cs.AI

摘要

当场景必须包含多个指定人物时,身份保持图像生成的可靠性会显著下降。除了保留每个身份外,模型还必须将每个参考对象绑定到不同的具体人物和位置,而训练时的身份损失函数必须在多个含噪声的预测人脸之间建立对应关系。我们提出WithEveryone,一个统一的框架,可生成包含多达十个参考身份的群像图像。WithEveryone将每个选定的身份注入为带寻址的令牌,预测结构化的身份-布局规划,并将该规划渲染为视觉条件。其关键目标函数——布局锚定身份损失(Layout-Grounded ID Loss)——利用标注的人脸区域直接监督目标身份,避免了基于嵌入的不稳定人脸匹配;身份表示强制(ID Representation Forcing)进一步在图像合成前对每个身份进行预测训练。在身份不相交的基准测试上,WithEveryone取得了最高的目标-上下文身份相似度,将人脸相似度从GPT-Image-2的0.462提升至0.499,同时将复制粘贴伪影从0.169降低至0.055。此外,该方法覆盖了97.3%的所需身份,重复率仅为2.8%。这些结果表明,显式的身份-布局锚定使身份保持生成能够扩展到更大规模的群组,而无需依赖直接参考人脸的复制。
English
Identity-preserving image generation becomes increasingly unreliable when a scene must contain many specified people. Beyond retaining each identity, the model must bind every reference to a distinct person and location, while training-time identity losses must establish correspondence among several noisy predicted faces. We introduce WithEveryone, a unified framework for generating group images up to ten reference identities. WithEveryone injects each selected identity as an addressed token, predicts a structured identity--layout plan, and renders the plan as a visual condition. Its key objective, Layout-Grounded ID Loss, uses annotated face regions to supervise the intended identities directly, avoiding unstable embedding-based face matching; ID Representation Forcing additionally trains a prediction for each identity before image synthesis. On an identity-disjoint benchmark, WithEveryone achieves the highest target-context identity similarity, improving face similarity from 0.462 for GPT-Image-2 to 0.499, while reducing copy-paste artifacts from 0.169 to 0.055. It further covers 97.3\% of the requested identities with a duplicate rate of only 2.8\%. These results show that explicit identity--layout grounding enables identity-preserving generation to scale to larger groups without relying on direct reference-face copying.