MultiRef-Compass: 다중 참조-오디오-비디오 생성의 포괄적 평가를 향하여
MultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video Generation
July 15, 2026
저자: Xiaohan Zhang, Yuqing Wen, Junlin Chen, Yuqi Tang, Yiting He, Lizhuo Shao, Weiming Zhu, Tengfei Liu, Yang Shi, Jialu Chen, Yuanxing Zhang, Huaxiong Li
cs.AI
초록
다중 참조-오디오-비디오(MR2AV) 생성은 여러 참조와 텍스트 명령어에 기반하여 일관된 오디오-비디오 콘텐츠를 생성하는 것을 목표로 한다. 기존 벤치마크는 주로 텍스트 기반 생성, 단일 참조 객체 보존, 또는 분리된 오디오-비디오 정렬에 초점을 맞추어, 새롭게 부상하는 MR2AV 설정은 대부분 탐구되지 않은 상태로 남아 있다. 이러한 설정과 비교할 때, MR2AV는 동기화된 시각 및 오디오 콘텐츠를 생성하면서 여러 참조에 대해 공동으로 추론하는 능력을 모델에 요구한다. 모델은 각 참조를 충실히 보존할 뿐만 아니라 여러 참조 개체를 올바르게 결합하고 구성하여 일관된 오디오-비디오 이벤트로 만들어야 한다. 이러한 격차를 해소하기 위해, 우리는 MR2AV 생성을 위한 통합 벤치마크인 MultiRef-Compass를 소개한다. 이 벤치마크는 확장 가능하고 제어 가능한 자산 구성 파이프라인을 통해 정성적으로 선별된 350개의 샘플로 구성되며, 다중 시점 객체 보존, 다중 개체 결합, 그리고 인간-객체-장면 구성을 포함한다. 해석 가능한 평가를 제공하기 위해 MultiRef-Compass는 14개의 하위 지표를 사용하여 기본 품질(Basic Quality), 참조 일관성(Reference Consistency), 오디오-비디오 일관성(Audio-Visual Consistency), 명령어 준수(Instruction Following)의 네 가지 차원으로 평가 프로토콜을 정의한다. MultiRef-Compass는 자동 평가 지표와 재판단 기능이 강화된 MLLM-as-a-Judge 프레임워크를 통합하여, 지각적 충실도와 참조 조건 구성 모두에 대한 확장 가능하고 감사 가능한 평가를 가능하게 한다. 8개의 대표적인 MR2AV 시스템에 대한 광범위한 실험은 여러 평가 차원에서 상당한 개선 여지가 있음을 보여주며, 포괄적인 벤치마크의 필요성을 강조하고 MultiRef-Compass를 향후 MR2AV 연구의 기반으로 자리매김한다.
English
Multi-reference-to-audio-video (MR2AV) generation aims to generate coherent audio-video content conditioned on multiple references and textual instructions. Existing benchmarks mainly focus on text-driven generation, single-reference subject preservation, or isolated audio-video alignment, leaving the emerging MR2AV setting largely unexplored. Compared with these settings, MR2AV requires models to jointly reason over multiple references while generating synchronized visual and audio content. Models must not only preserve each reference faithfully but also correctly bind and compose multiple referenced entities into coherent audio-visual events. To address this gap, we introduce MultiRef-Compass, a unified benchmark for MR2AV generation. It comprises 350 carefully curated samples constructed through a scalable and controllable asset-composition pipeline, covering multi-view subject preservation, multi-entity binding, and human-object-scene composition. To provide interpretable assessment, MultiRef-Compass defines an evaluation protocol with four dimensions: Basic Quality, Reference Consistency, Audio-Visual Consistency, and Instruction Following, using 14 sub-metrics. MultiRef-Compass integrates automatic metrics with a rejudging-enhanced MLLM-as-a-Judge framework, enabling scalable and auditable evaluation of both perceptual fidelity and reference-conditioned composition. Extensive experiments on eight representative MR2AV systems reveal substantial room for improvement across multiple evaluation dimensions, underscoring the need for a comprehensive benchmark and positioning MultiRef-Compass as a foundation for future MR2AV research.