MPIE-Bench:解剖学合理的多人交互编辑基准评测
MPIE-Bench: Benchmarking Anatomically Plausible Multi-Person Interaction Editing
July 30, 2026
作者: Jiajia Lin, Mingxuan Du, Tuowen Zhou, Benfeng Xu, Hongtao Xie
cs.AI
摘要
文本到图像及个性化编辑模型如今能够轻松合成高保真的单主体图像。然而,将多个具名人物置于共享接触动作(如拥抱、背负或扭打)中,仍暴露出重大缺陷:肢体融合、多余肢体出现以及身体相互穿透。现有评估大多忽视了这些解剖学和几何学问题,而“VLM即评判者”式的检查清单在交互维度上往往趋于饱和,尽管这些错误对人类而言显而易见。我们提出MPIE-Bench,一个包含2,500个样本的基准数据集,其样本源自视频挖掘的编辑三元组,涵盖405个场景、14种交互类别和四档接触密度(C0-C3)。同时,我们提出MPIE-Eval,其新增的两个评估轴通过冻结的公开多人体网格重建方法对接触时刻的几何质量进行评分。“解剖学”轴考察是否每一处类人体质量都能由一组完整重建的身体所解释,“交互”轴则考察这些身体之间的穿透程度和表面距离是否与指令所要求的接触相匹配。在十个编辑模型上的实验表明,网格解剖学得分在两种不同模型上最高仅为0.65,网格交互得分最高为0.72,因此没有任何单一编辑模型能在两个维度上均表现优异,而VLM检查清单对相同图像的评分却高于0.95。一项由五名评分员参与的研究证实,这两个评估轴均比零样本VLM评判者更贴近人类判断,并且在对所有权重和阈值进行消融后,排名结果依然保持稳定。
English
Text-to-image and personalized editing models now synthesize high-fidelity single-subject images with ease. Yet placing multiple named people into shared contact actions such as embrace, carry, or grapple still exposes major failures: fused limbs, invented extremities, and interpenetrating bodies. Existing evaluations largely overlook these anatomical and geometric issues, and VLM-as-a-judge checklists often saturate on Interaction while the errors remain obvious to humans. We introduce MPIE-Bench, a 2,500-sample benchmark of video-mined editing triplets spanning 405 scenes, 14 interaction categories, and four contact densities (C0-C3). We also propose MPIE-Eval, whose two new axes score contact-time geometry from a frozen public multi-person mesh reconstruction. Anatomy asks whether every human-like mass is explained by a complete set of reconstructed bodies, and Interaction asks whether the penetration and surface distance between those bodies match the contact the instruction asked for. Across ten editors, mesh Anatomy tops out at 0.65 and mesh Interaction at 0.72 on two different models, so no single editor is strong on both, while VLM checklists rate the same images above 0.95. A five-rater study confirms that both axes track human judgement more closely than a zero-shot VLM judge, and the rankings hold under ablation of every weight and threshold.