ChatPaper.aiChatPaper

TRACE-Bench:多參考影像生成之分解與診斷

TRACE-Bench: Decomposing and Diagnosing Multi-Reference Image Generation

August 17, 2026
作者: Haoran Wang, Chaofan Ma, Ran Yi, Lizhuang Ma
cs.AI

摘要

儘管近年來統一多模態模型在多參考圖像生成方面取得了進展,現有的基準仍圍繞預定義任務類型(例如「主題組合」)進行組織,而這種方式並不適合此類組合設定,導致覆蓋範圍零散、複雜度不可控且診斷價值有限。鑒於多樣化的多參考任務共享一組共同的原子操作,我們採取以能力為導向的視角,並形式化定義四個算子:錨定(Anchor, f)、解耦(Disentangle, g)、應用(Apply, ⊕)與組合(Compose, C)。任何多參考提示詞皆可表示為這些算子上的組合公式,其結構複雜度以算子槽位數量進行量化。基於此形式化框架,我們建構了 TRACE-Bench,包含約 1,600 個評估案例,槽位數涵蓋 1 至 8,由 631 個公式模板與約 4,000 張涵蓋多種藝術風格及真實世界主體的參考圖像所構成。公式結構直接驅動了算子對齊的評估協議,以進行按能力評分,並透過診斷樹分析實現遞迴失敗定位。對 9 個領先模型的評估揭示出整體評分無法察覺的洞見:主要瓶頸在於解耦(g)與屬性綁定(⊕),而非場景級組合(C);即使是最佳模型,其屬性保真度得分也僅為 0.74。專案頁面:https://amuseum-whr.github.io/TraceBench
English
Despite recent advances in unified multimodal models for multi-reference image generation, existing benchmarks remain organized around predefined task types (e.g., "subject composition"), which are ill-suited to this combinatorial setting and lead to fragmented coverage, uncontrolled complexity, and little diagnostic value. Recognizing that diverse multi-reference tasks share a common set of atomic operations, we adopt a capability-oriented perspective and formalize four operators: Anchor (f), Disentangle (g), Apply (oplus), and Compose (C). Any multi-reference prompt can then be represented as a compositional formula over these operators, whose structural complexity is quantified by the number of operator slots. Building on this formulation, we construct TRACE-Bench, comprising approximately 1,600 evaluation cases across slot counts 1--8, built from 631 formula templates and around 4,000 reference images spanning diverse artistic styles and real-world subjects. The formula structure directly drives an operator-aligned evaluation protocol for per-capability scoring and a diagnostic tree analysis for recursive failure localization. Evaluating 9 leading models reveals insights invisible to holistic scoring: the primary bottleneck lies in disentanglement (g) and attribute binding (oplus) rather than scene-level composition (C), with even the best model scoring only 0.74 on attribute fidelity. Project page: https://amuseum-whr.github.io/TraceBench