MultiRef-Compass:邁向多參考音視頻生成之全面評估
MultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video Generation
July 15, 2026
作者: Xiaohan Zhang, Yuqing Wen, Junlin Chen, Yuqi Tang, Yiting He, Lizhuo Shao, Weiming Zhu, Tengfei Liu, Yang Shi, Jialu Chen, Yuanxing Zhang, Huaxiong Li
cs.AI
摘要
多參考音視頻生成(MR2AV)旨在基於多個參考與文字指令生成連貫的音視頻內容。現有基準測試主要聚焦於文字驅動生成、單一參考主體保留或獨立的音視頻對齊,使得新興的MR2AV設定在很大程度上未獲探索。與這些設定相比,MR2AV要求模型在生成同步的視覺與音頻內容時,對多個參考進行聯合推理。模型不僅須忠實保留每個參考,還須正確綁定並組合多個參考實體,形成連貫的音視頻事件。為填補此缺口,我們提出MultiRef-Compass,一個用於MR2AV生成的統一基準。該基準包含350個精心策劃的樣本,經由可擴展且可控的資產組合管道構建,涵蓋多視角主體保留、多實體綁定以及人物-物體-場景組合。為提供可解釋的評估,MultiRef-Compass定義了一個包含四個維度的評估協議:基本品質、參考一致性、音視頻一致性與指令遵循,並使用14項子指標。MultiRef-Compass將自動化指標與重新評判增強的MLLM作為評判框架相結合,實現對感知保真度與參考條件下組合的可擴展且可審計的評估。在八個具代表性的MR2AV系統上進行的廣泛實驗顯示,多個評估維度仍有顯著改善空間,凸顯全面基準測試的必要性,並將MultiRef-Compass定位為未來MR2AV研究的基礎。
English
Multi-reference-to-audio-video (MR2AV) generation aims to generate coherent audio-video content conditioned on multiple references and textual instructions. Existing benchmarks mainly focus on text-driven generation, single-reference subject preservation, or isolated audio-video alignment, leaving the emerging MR2AV setting largely unexplored. Compared with these settings, MR2AV requires models to jointly reason over multiple references while generating synchronized visual and audio content. Models must not only preserve each reference faithfully but also correctly bind and compose multiple referenced entities into coherent audio-visual events. To address this gap, we introduce MultiRef-Compass, a unified benchmark for MR2AV generation. It comprises 350 carefully curated samples constructed through a scalable and controllable asset-composition pipeline, covering multi-view subject preservation, multi-entity binding, and human-object-scene composition. To provide interpretable assessment, MultiRef-Compass defines an evaluation protocol with four dimensions: Basic Quality, Reference Consistency, Audio-Visual Consistency, and Instruction Following, using 14 sub-metrics. MultiRef-Compass integrates automatic metrics with a rejudging-enhanced MLLM-as-a-Judge framework, enabling scalable and auditable evaluation of both perceptual fidelity and reference-conditioned composition. Extensive experiments on eight representative MR2AV systems reveal substantial room for improvement across multiple evaluation dimensions, underscoring the need for a comprehensive benchmark and positioning MultiRef-Compass as a foundation for future MR2AV research.