ChatPaper.aiChatPaper

MultiRef-Compass:面向多参考音视频生成的综合评估

MultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video Generation

July 15, 2026
作者: Xiaohan Zhang, Yuqing Wen, Junlin Chen, Yuqi Tang, Yiting He, Lizhuo Shao, Weiming Zhu, Tengfei Liu, Yang Shi, Jialu Chen, Yuanxing Zhang, Huaxiong Li
cs.AI

摘要

多参考到音视频生成(Multi-reference-to-audio-video, MR2AV)旨在基于多个参考样本与文本指令生成连贯的音频-视频内容。现有基准主要关注文本驱动生成、单参考主体保持或孤立的音视频对齐,而新兴的MR2AV设置尚未得到充分探索。与这些设置相比,MR2AV要求模型在生成同步视觉与音频内容的同时,对多个参考样本进行联合推理。模型不仅需要忠实保持每个参考样本的特征,还需正确绑定与组合多个被参考实体,以形成连贯的视听事件。为填补这一空白,我们提出MultiRef-Compass——用于MR2AV生成的统一基准。该基准包含350个精心构建的样本,通过可扩展且可控的资产组合流水线生成,覆盖多视角主体保持、多实体绑定以及人-物-场景组合。为提供可解释的评估,MultiRef-Compass定义了包含四个维度的评估协议:基础质量、参考一致性、视听一致性及指令遵循,并采用14项子指标。MultiRef-Compass将自动指标与基于重判增强的多模态大语言模型作为评判者(MLLM-as-a-Judge)框架相结合,实现了对感知保真度与参考条件组合的可扩展、可审计评估。在八个代表性MR2AV系统上的广泛实验揭示,多个评估维度仍存在显著改进空间,凸显了全面基准的必要性,并使MultiRef-Compass成为未来MR2AV研究的基础平台。
English
Multi-reference-to-audio-video (MR2AV) generation aims to generate coherent audio-video content conditioned on multiple references and textual instructions. Existing benchmarks mainly focus on text-driven generation, single-reference subject preservation, or isolated audio-video alignment, leaving the emerging MR2AV setting largely unexplored. Compared with these settings, MR2AV requires models to jointly reason over multiple references while generating synchronized visual and audio content. Models must not only preserve each reference faithfully but also correctly bind and compose multiple referenced entities into coherent audio-visual events. To address this gap, we introduce MultiRef-Compass, a unified benchmark for MR2AV generation. It comprises 350 carefully curated samples constructed through a scalable and controllable asset-composition pipeline, covering multi-view subject preservation, multi-entity binding, and human-object-scene composition. To provide interpretable assessment, MultiRef-Compass defines an evaluation protocol with four dimensions: Basic Quality, Reference Consistency, Audio-Visual Consistency, and Instruction Following, using 14 sub-metrics. MultiRef-Compass integrates automatic metrics with a rejudging-enhanced MLLM-as-a-Judge framework, enabling scalable and auditable evaluation of both perceptual fidelity and reference-conditioned composition. Extensive experiments on eight representative MR2AV systems reveal substantial room for improvement across multiple evaluation dimensions, underscoring the need for a comprehensive benchmark and positioning MultiRef-Compass as a foundation for future MR2AV research.