TRACE-Bench: 다중 참조 이미지 생성의 분해 및 진단
TRACE-Bench: Decomposing and Diagnosing Multi-Reference Image Generation
August 17, 2026
저자: Haoran Wang, Chaofan Ma, Ran Yi, Lizhuang Ma
cs.AI
초록
최근 다중 참조 이미지 생성을 위한 통합 멀티모달 모델의 발전에도 불구하고, 기존 벤치마크는 사전 정의된 과업 유형(예: '객체 구성')을 중심으로 구성되어 있어, 이러한 조합적 설정에 부적합하며 단편적 커버리지, 통제되지 않은 복잡성, 낮은 진단 가치를 초래한다. 다양한 다중 참조 과업들이 공통된 원자적 연산들을 공유한다는 점에 착안하여, 우리는 능력 중심 관점을 채택하고 네 가지 연산자, 즉 Anchor (f), Disentangle (g), Apply (⊕), Compose (C)를 정식화한다. 그러면 모든 다중 참조 프롬프트는 이 연산자들에 대한 조합 공식으로 표현될 수 있으며, 그 구조적 복잡성은 연산자 슬롯 수로 정량화된다. 이러한 정식화를 바탕으로 우리는 TRACE-Bench를 구축한다. TRACE-Bench는 슬롯 수 1~8에 걸친 약 1,600개의 평가 사례로 구성되며, 631개의 공식 템플릿과 다양한 예술적 스타일 및 실제 세계 객체를 아우르는 약 4,000개의 참조 이미지로 구축된다. 공식 구조는 능력별 점수 산출을 위한 연산자에 맞춘 평가 프로토콜과 재귀적 실패 위치 파악을 위한 진단 트리 분석을 직접 유도한다. 9개의 주요 모델을 평가한 결과, 총체적 점수로는 드러나지 않는 통찰이 얻어진다. 주요 병목은 장면 수준 구성(C)이 아니라 분리(g)와 속성 바인딩(⊕)에 있으며, 최고 성능 모델조차 속성 충실도에서 0.74에 그친다. 프로젝트 페이지: https://amuseum-whr.github.io/TraceBench
English
Despite recent advances in unified multimodal models for multi-reference image generation, existing benchmarks remain organized around predefined task types (e.g., "subject composition"), which are ill-suited to this combinatorial setting and lead to fragmented coverage, uncontrolled complexity, and little diagnostic value. Recognizing that diverse multi-reference tasks share a common set of atomic operations, we adopt a capability-oriented perspective and formalize four operators: Anchor (f), Disentangle (g), Apply (oplus), and Compose (C). Any multi-reference prompt can then be represented as a compositional formula over these operators, whose structural complexity is quantified by the number of operator slots. Building on this formulation, we construct TRACE-Bench, comprising approximately 1,600 evaluation cases across slot counts 1--8, built from 631 formula templates and around 4,000 reference images spanning diverse artistic styles and real-world subjects. The formula structure directly drives an operator-aligned evaluation protocol for per-capability scoring and a diagnostic tree analysis for recursive failure localization. Evaluating 9 leading models reveals insights invisible to holistic scoring: the primary bottleneck lies in disentanglement (g) and attribute binding (oplus) rather than scene-level composition (C), with even the best model scoring only 0.74 on attribute fidelity. Project page: https://amuseum-whr.github.io/TraceBench