ChatPaper.aiChatPaper

MPIE-Bench: 해부학적으로 타당한 다중 인물 상호작용 편집을 위한 벤치마크

MPIE-Bench: Benchmarking Anatomically Plausible Multi-Person Interaction Editing

July 30, 2026
저자: Jiajia Lin, Mingxuan Du, Tuowen Zhou, Benfeng Xu, Hongtao Xie
cs.AI

초록

텍스트-이미지 및 개인화 편집 모델은 이제 단일 대상 이미지를 높은 충실도로 손쉽게 합성한다. 그러나 포옹, 들기, 또는 맞붙기와 같은 공동 접촉 동작에서 여러 명의 지명된 인물을 배치하는 작업은 여전히 심각한 오류를 노출한다: 융합된 사지, 생성된 부속지, 그리고 상호 관통하는 신체가 그것이다. 기존 평가는 이러한 해부학적 및 기하학적 문제를 대체로 간과하며, VLM-심판 체크리스트는 상호작용 항목에서 종종 포화 상태에 이르는 반면 오류는 인간에게 명백하게 남아 있다. 우리는 405개 장면, 14개 상호작용 범주, 4가지 접촉 밀도(C0-C3)에 걸친 2,500개 샘플의 비디오 기반 편집 트리플릿 벤치마크인 MPIE-Bench를 소개한다. 또한 우리는 동결된 공개 다중 인물 메시 재구성에서 접촉 시간 기하 구조를 평가하는 두 개의 새로운 축을 갖춘 MPIE-Eval을 제안한다. 해부학(Anatomy) 축은 모든 인간형 질량이 완전한 재구성 신체 집합으로 설명되는지 여부를 묻고, 상호작용(Interaction) 축은 해당 신체들 사이의 관통 정도와 표면 거리가 지시문이 요청한 접촉과 일치하는지 여부를 묻는다. 10개 편집 모델에 걸쳐, 메시 기반 해부학 점수는 서로 다른 두 모델에서 최대 0.65, 메시 기반 상호작용 점수는 최대 0.72에 그쳐, 두 축 모두에서 강한 성능을 보이는 단일 편집 모델은 없었다. 반면 VLM 체크리스트는 동일한 이미지들을 0.95 이상으로 평가했다. 5인의 평가자 연구는 두 축 모두 제로샷 VLM 심판보다 인간 판단에 더 밀접하게 부합함을 확인했으며, 이러한 순위는 모든 가중치와 임계값의 제거 실험에서도 유지된다.
English
Text-to-image and personalized editing models now synthesize high-fidelity single-subject images with ease. Yet placing multiple named people into shared contact actions such as embrace, carry, or grapple still exposes major failures: fused limbs, invented extremities, and interpenetrating bodies. Existing evaluations largely overlook these anatomical and geometric issues, and VLM-as-a-judge checklists often saturate on Interaction while the errors remain obvious to humans. We introduce MPIE-Bench, a 2,500-sample benchmark of video-mined editing triplets spanning 405 scenes, 14 interaction categories, and four contact densities (C0-C3). We also propose MPIE-Eval, whose two new axes score contact-time geometry from a frozen public multi-person mesh reconstruction. Anatomy asks whether every human-like mass is explained by a complete set of reconstructed bodies, and Interaction asks whether the penetration and surface distance between those bodies match the contact the instruction asked for. Across ten editors, mesh Anatomy tops out at 0.65 and mesh Interaction at 0.72 on two different models, so no single editor is strong on both, while VLM checklists rate the same images above 0.95. A five-rater study confirms that both axes track human judgement more closely than a zero-shot VLM judge, and the rankings hold under ablation of every weight and threshold.