ChatPaper.aiChatPaper

TRACE-Bench:分解与诊断多参考图像生成

TRACE-Bench: Decomposing and Diagnosing Multi-Reference Image Generation

August 17, 2026
作者: Haoran Wang, Chaofan Ma, Ran Yi, Lizhuang Ma
cs.AI

摘要

尽管近年来面向多参考图像生成的统一多模态模型取得了显著进展,但现有基准测试仍围绕预定义任务类型(如“主体组合”)来组织,难以适应这种组合性场景,导致覆盖碎片化、复杂度不可控,且诊断价值有限。我们认识到多样化的多参考任务共享一组原子操作,因此采用能力导向的视角,形式化了四个算子:锚定(f)、解耦(g)、应用(⊕)和组合(C)。由此,任何多参考提示词都可以表示为基于这些算子的组合公式,其结构复杂度通过算子槽位数量来量化。基于这一形式化框架,我们构建了 TRACE-Bench,包含约 1,600 个评测用例,覆盖 1 至 8 个槽位数量,由 631 个公式模板和约 4,000 张涵盖多样艺术风格与真实世界主体的参考图像构建而成。公式结构直接驱动了用于按能力评分的算子对齐评估协议,以及用于递归失败定位的诊断树分析。对 9 个领先模型的评估揭示了整体评分无法呈现的洞见:主要瓶颈在于解耦(g)和属性绑定(⊕),而非场景级组合(C);即使最佳模型在属性保真度上也仅获得 0.74 分。项目主页:https://amuseum-whr.github.io/TraceBench
English
Despite recent advances in unified multimodal models for multi-reference image generation, existing benchmarks remain organized around predefined task types (e.g., "subject composition"), which are ill-suited to this combinatorial setting and lead to fragmented coverage, uncontrolled complexity, and little diagnostic value. Recognizing that diverse multi-reference tasks share a common set of atomic operations, we adopt a capability-oriented perspective and formalize four operators: Anchor (f), Disentangle (g), Apply (oplus), and Compose (C). Any multi-reference prompt can then be represented as a compositional formula over these operators, whose structural complexity is quantified by the number of operator slots. Building on this formulation, we construct TRACE-Bench, comprising approximately 1,600 evaluation cases across slot counts 1--8, built from 631 formula templates and around 4,000 reference images spanning diverse artistic styles and real-world subjects. The formula structure directly drives an operator-aligned evaluation protocol for per-capability scoring and a diagnostic tree analysis for recursive failure localization. Evaluating 9 leading models reveals insights invisible to holistic scoring: the primary bottleneck lies in disentanglement (g) and attribute binding (oplus) rather than scene-level composition (C), with even the best model scoring only 0.74 on attribute fidelity. Project page: https://amuseum-whr.github.io/TraceBench