WithEveryone: 그룹 이미지 생성을 위한 통합 계획 및 정체성 기반
WithEveryone: Unified Planning and Identity Grounding for Group Image Generation
August 20, 2026
저자: Hengyuan Xu, Qixun Wang, Yiji Cheng, Miles Yang, Zhao Zhong, Wei Cheng, Xingjun Ma, Yu-gang Jiang
cs.AI
초록
정체성 보존 이미지 생성은 장면에 많은 특정 인물이 포함되어야 할 때 점점 더 신뢰할 수 없게 된다. 각 정체성을 유지하는 것 외에도, 모델은 모든 참조를 서로 다른 인물 및 위치에 연결해야 하며, 학습 시점의 정체성 손실은 여러 노이즈가 있는 예측 얼굴들 사이의 대응 관계를 설정해야 한다. 우리는 최대 10개의 참조 정체성을 가진 그룹 이미지를 생성하는 통합 프레임워크인 WithEveryone을 소개한다. WithEveryone은 선택된 각 정체성을 주소 지정된 토큰으로 주입하고, 구조화된 정체성-레이아웃 계획을 예측하며, 이 계획을 시각적 조건으로 렌더링한다. 핵심 목적 함수인 Layout-Grounded ID Loss는 주석이 달린 얼굴 영역을 사용하여 의도된 정체성을 직접 지도함으로써 불안정한 임베딩 기반 얼굴 매칭을 피한다. ID Representation Forcing은 또한 이미지 합성 전에 각 정체성에 대한 예측을 추가로 훈련한다. 정체성이 분리된 벤치마크에서 WithEveryone은 가장 높은 대상-맥락 정체성 유사도를 달성하여 GPT-Image-2의 0.462에서 0.499로 얼굴 유사도를 향상시키고, 복사-붙여넣기 아티팩트를 0.169에서 0.055로 줄인다. 또한 요청된 정체성의 97.3%를 포함하며 중복률은 2.8%에 불과하다. 이러한 결과는 명시적인 정체성-레이아웃 접지가 직접적인 참조 얼굴 복사에 의존하지 않고도 정체성 보존 생성을 더 큰 그룹으로 확장할 수 있게 함을 보여준다.
English
Identity-preserving image generation becomes increasingly unreliable when a scene must contain many specified people. Beyond retaining each identity, the model must bind every reference to a distinct person and location, while training-time identity losses must establish correspondence among several noisy predicted faces. We introduce WithEveryone, a unified framework for generating group images up to ten reference identities. WithEveryone injects each selected identity as an addressed token, predicts a structured identity--layout plan, and renders the plan as a visual condition. Its key objective, Layout-Grounded ID Loss, uses annotated face regions to supervise the intended identities directly, avoiding unstable embedding-based face matching; ID Representation Forcing additionally trains a prediction for each identity before image synthesis. On an identity-disjoint benchmark, WithEveryone achieves the highest target-context identity similarity, improving face similarity from 0.462 for GPT-Image-2 to 0.499, while reducing copy-paste artifacts from 0.169 to 0.055. It further covers 97.3\% of the requested identities with a duplicate rate of only 2.8\%. These results show that explicit identity--layout grounding enables identity-preserving generation to scale to larger groups without relying on direct reference-face copying.