MPIE-Bench: 解剖学的に妥当な複数人インタラクション編集のベンチマーク評価
MPIE-Bench: Benchmarking Anatomically Plausible Multi-Person Interaction Editing
July 30, 2026
著者: Jiajia Lin, Mingxuan Du, Tuowen Zhou, Benfeng Xu, Hongtao Xie
cs.AI
要旨
テキストから画像を生成するモデルやパーソナライズ編集モデルは、現在では高忠実度の単一被写体画像を容易に合成できる。しかし、抱擁、運搬、格闘といった複数の名前付き人物が共有する接触動作を生成する場合、四肢の融合、幻の四肢、身体同士の貫入といった重大な失敗が依然として発生する。既存の評価手法はこれらの解剖学的・幾何学的問題をほとんど考慮しておらず、VLMを判定者とするチェックリストはInteraction軸において飽和しがちである一方、人間にはこれらの誤りが明らかである。我々は、405シーン、14のインタラクションカテゴリ、4つの接触密度(C0〜C3)にわたる2,500サンプルのビデオマイニング編集トリプレットからなるベンチマークMPIE-Benchを導入する。さらに、固定した公開の多人数メッシュ再構成から接触時の幾何形状を評価する2つの新しい軸を備えたMPIE-Evalを提案する。Anatomy軸は、すべての人体状の塊が完全な再構成身体群によって説明されるかを問い、Interaction軸は、それらの身体間の貫入量と表面距離が指示された接触と一致するかを問う。10種類のエディタの評価において、メッシュベースのAnatomyは最大でも0.65、メッシュベースのInteractionは最大でも0.72に留まり、両方に優れた単一エディタは存在しない一方、VLMチェックリストは同じ画像を0.95以上と評価する。5人の評価者による研究は、両軸がゼロショットVLM判定者よりも人間の判断をより忠実に反映することを確認し、そのランキングはすべての重みと閾値のアブレーションでも維持される。
English
Text-to-image and personalized editing models now synthesize high-fidelity single-subject images with ease. Yet placing multiple named people into shared contact actions such as embrace, carry, or grapple still exposes major failures: fused limbs, invented extremities, and interpenetrating bodies. Existing evaluations largely overlook these anatomical and geometric issues, and VLM-as-a-judge checklists often saturate on Interaction while the errors remain obvious to humans. We introduce MPIE-Bench, a 2,500-sample benchmark of video-mined editing triplets spanning 405 scenes, 14 interaction categories, and four contact densities (C0-C3). We also propose MPIE-Eval, whose two new axes score contact-time geometry from a frozen public multi-person mesh reconstruction. Anatomy asks whether every human-like mass is explained by a complete set of reconstructed bodies, and Interaction asks whether the penetration and surface distance between those bodies match the contact the instruction asked for. Across ten editors, mesh Anatomy tops out at 0.65 and mesh Interaction at 0.72 on two different models, so no single editor is strong on both, while VLM checklists rate the same images above 0.95. A five-rater study confirms that both axes track human judgement more closely than a zero-shot VLM judge, and the rankings hold under ablation of every weight and threshold.