TRACE-Bench: マルチ参照画像生成の分解と診断
TRACE-Bench: Decomposing and Diagnosing Multi-Reference Image Generation
August 17, 2026
著者: Haoran Wang, Chaofan Ma, Ran Yi, Lizhuang Ma
cs.AI
要旨
複数参照画像生成のための統合マルチモーダルモデルは近年進歩しているものの、既存のベンチマークは(例えば「被写体合成」のような)事前定義されたタスク種別を中心に構成されたままであり、このような組合せ的設定には適しておらず、断片的なカバレッジ、制御されていない複雑性、そして診断的価値の乏しさをもたらしている。
多様な複数参照タスクが共通の原子的操作を共有していることを認識し、我々は能力指向の視点を採用して、アンカー (f)、分離 (g)、適用 (⊕)、合成 (C) の4つの演算子を形式化する。これにより、任意の複数参照プロンプトは、これらの演算子からなる合成的な式として表現でき、その構造的複雑さは演算子スロットの数によって定量化される。
この定式化に基づき、我々は TRACE-Bench を構築する。これは、スロット数1〜8にわたる約1,600件の評価ケースからなり、631個の式テンプレートと、多様な芸術スタイルおよび実世界の被写体を網羅する約4,000枚の参照画像から構築されている。
この式構造は、能力ごとのスコアリングのための演算子整合型評価プロトコルと、再帰的な障害位置特定のための診断ツリー分析を直接駆動する。
9つの主要モデルを評価したところ、全体的なスコアリングでは見えない洞察が明らかになった。すなわち、主要なボトルネックはシーンレベルの合成 (C) ではなく、分離 (g) と属性結合 (⊕) にあり、最良のモデルでも属性忠実度は0.74に留まる。
プロジェクトページ: https://amuseum-whr.github.io/TraceBench
English
Despite recent advances in unified multimodal models for multi-reference image generation, existing benchmarks remain organized around predefined task types (e.g., "subject composition"), which are ill-suited to this combinatorial setting and lead to fragmented coverage, uncontrolled complexity, and little diagnostic value. Recognizing that diverse multi-reference tasks share a common set of atomic operations, we adopt a capability-oriented perspective and formalize four operators: Anchor (f), Disentangle (g), Apply (oplus), and Compose (C). Any multi-reference prompt can then be represented as a compositional formula over these operators, whose structural complexity is quantified by the number of operator slots. Building on this formulation, we construct TRACE-Bench, comprising approximately 1,600 evaluation cases across slot counts 1--8, built from 631 formula templates and around 4,000 reference images spanning diverse artistic styles and real-world subjects. The formula structure directly drives an operator-aligned evaluation protocol for per-capability scoring and a diagnostic tree analysis for recursive failure localization. Evaluating 9 leading models reveals insights invisible to holistic scoring: the primary bottleneck lies in disentanglement (g) and attribute binding (oplus) rather than scene-level composition (C), with even the best model scoring only 0.74 on attribute fidelity. Project page: https://amuseum-whr.github.io/TraceBench