ChatPaper.aiChatPaper

MultiRef-Compass: マルチリファレンスからの音声・映像生成の包括的評価に向けて

MultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video Generation

July 15, 2026
著者: Xiaohan Zhang, Yuqing Wen, Junlin Chen, Yuqi Tang, Yiting He, Lizhuo Shao, Weiming Zhu, Tengfei Liu, Yang Shi, Jialu Chen, Yuanxing Zhang, Huaxiong Li
cs.AI

要旨

マルチ参照から音声・映像(MR2AV)生成は、複数の参照とテキスト指示に基づいて整合性のある音声・映像コンテンツを生成することを目的とする。既存のベンチマークは主にテキスト駆動生成、単一参照の被写体保存、または独立した音声・映像の整合性に焦点を当てており、新たに登場したMR2AV設定はほとんど未開拓のままである。これらの設定と比較すると、MR2AVはモデルに、同期した視覚・音声コンテンツを生成する際に複数の参照を共同で推論することを要求する。モデルは各参照を忠実に保存するだけでなく、複数の参照エンティティを正しく結びつけ、整合性のある音声・映像イベントに合成する必要がある。このギャップに対処するため、我々はMR2AV生成のための統一ベンチマークであるMultiRef-Compassを導入する。これは、スケーラブルで制御可能なアセット合成パイプラインを通じて構築された350の厳選されたサンプルから成り、マルチビュー被写体保存、マルチエンティティ結合、人間-物体-シーン合成をカバーする。解釈可能な評価を提供するため、MultiRef-Compassは14のサブ指標を使用して、基本品質、参照一貫性、音声・映像一貫性、指示追従の4次元からなる評価プロトコルを定義する。MultiRef-Compassは、自動指標と再判定強化型MLLM-as-a-Judgeフレームワークを統合し、知覚的忠実性と参照条件付き合成の両方のスケーラブルで監査可能な評価を可能にする。8つの代表的なMR2AVシステムに対する広範な実験は、複数の評価次元において改善の余地が大きいことを明らかにし、包括的なベンチマークの必要性を強調し、MultiRef-Compassを将来のMR2AV研究の基盤として位置づける。
English
Multi-reference-to-audio-video (MR2AV) generation aims to generate coherent audio-video content conditioned on multiple references and textual instructions. Existing benchmarks mainly focus on text-driven generation, single-reference subject preservation, or isolated audio-video alignment, leaving the emerging MR2AV setting largely unexplored. Compared with these settings, MR2AV requires models to jointly reason over multiple references while generating synchronized visual and audio content. Models must not only preserve each reference faithfully but also correctly bind and compose multiple referenced entities into coherent audio-visual events. To address this gap, we introduce MultiRef-Compass, a unified benchmark for MR2AV generation. It comprises 350 carefully curated samples constructed through a scalable and controllable asset-composition pipeline, covering multi-view subject preservation, multi-entity binding, and human-object-scene composition. To provide interpretable assessment, MultiRef-Compass defines an evaluation protocol with four dimensions: Basic Quality, Reference Consistency, Audio-Visual Consistency, and Instruction Following, using 14 sub-metrics. MultiRef-Compass integrates automatic metrics with a rejudging-enhanced MLLM-as-a-Judge framework, enabling scalable and auditable evaluation of both perceptual fidelity and reference-conditioned composition. Extensive experiments on eight representative MR2AV systems reveal substantial room for improvement across multiple evaluation dimensions, underscoring the need for a comprehensive benchmark and positioning MultiRef-Compass as a foundation for future MR2AV research.