MPIE-Bench:解剖學上合理的多人互動編輯基準評測
MPIE-Bench: Benchmarking Anatomically Plausible Multi-Person Interaction Editing
July 30, 2026
作者: Jiajia Lin, Mingxuan Du, Tuowen Zhou, Benfeng Xu, Hongtao Xie
cs.AI
摘要
文字生成影像與個人化編輯模型現在能輕鬆合成高保真度的單一主體影像。然而,要將多位具名人物置入如擁抱、抱起或扭打等共享接觸動作中,仍會暴露重大缺陷:肢體融合、虛構的四肢,以及身體相互穿透。現有評估大多忽略這些解剖學與幾何學問題,而以視覺語言模型(VLM)作為評審的檢查清單在「互動」(Interaction)指標上往往達到飽和,但這些錯誤對人類而言仍顯而易見。我們提出MPIE-Bench,一個包含2,500個樣本的基準測試,取材自影片探勘的編輯三元組,涵蓋405個場景、14個互動類別與四種接觸密度(C0-C3)。我們亦提出MPIE-Eval,其兩個新評估軸向利用凍結的公開多人網格重建來評分接觸時的幾何關係。「解剖學」(Anatomy)軸向探討每一塊類人質量是否皆能被一組完整的重建身體所解釋,而「互動」(Interaction)軸向則探討這些身體之間的穿透量與表面距離是否符合指令所要求的接觸程度。在十個編輯器的測試中,兩個不同模型上的網格解剖學指標最高僅達0.65,網格互動指標最高僅達0.72,因此沒有任何單一編輯器能在兩個軸向上皆表現優異,而VLM檢查清單對相同影像的評分卻高於0.95。一項五位評分員的研究證實,這兩個軸向比零樣本VLM評審更貼近人類判斷,且在對每個權重與閾值進行消融測試後,排名依然成立。
English
Text-to-image and personalized editing models now synthesize high-fidelity single-subject images with ease. Yet placing multiple named people into shared contact actions such as embrace, carry, or grapple still exposes major failures: fused limbs, invented extremities, and interpenetrating bodies. Existing evaluations largely overlook these anatomical and geometric issues, and VLM-as-a-judge checklists often saturate on Interaction while the errors remain obvious to humans. We introduce MPIE-Bench, a 2,500-sample benchmark of video-mined editing triplets spanning 405 scenes, 14 interaction categories, and four contact densities (C0-C3). We also propose MPIE-Eval, whose two new axes score contact-time geometry from a frozen public multi-person mesh reconstruction. Anatomy asks whether every human-like mass is explained by a complete set of reconstructed bodies, and Interaction asks whether the penetration and surface distance between those bodies match the contact the instruction asked for. Across ten editors, mesh Anatomy tops out at 0.65 and mesh Interaction at 0.72 on two different models, so no single editor is strong on both, while VLM checklists rate the same images above 0.95. A five-rater study confirms that both axes track human judgement more closely than a zero-shot VLM judge, and the rankings hold under ablation of every weight and threshold.