AVE-Compass: 音声・映像編集能力の総合的評価に向けて
AVE-Compass: Towards Holistic Evaluation for Audio-Video Editing Abilities
July 17, 2026
著者: Yuqing Wen, Yukai Huang, Qianqian Xie, Jiangtao Wu, Yibin Lin, Yikai Gu, Jialu Chen, Yuanxing Zhang, Jiaheng Liu
cs.AI
要旨
指示ベースのビデオ編集は急速に進歩しているが、実世界のビデオには音声信号と視覚信号が密接に結合されており、一方のモダリティを編集する際には、しばしば他方のモダリティの協調的な変更が必要となる。既存のベンチマークは主に無音クリップに対する視覚的変換や、独立した音声編集を評価しており、複雑な音声・視覚編集とクロスモーダル一貫性は未探索のままである。我々はAVE-Compassを紹介する。これは、145本の厳選されたソースビデオ、196件の音声・視覚が結合された編集指示、2,688項目の詳細なチェックリスト項目からなる包括的なベンチマークである。本ベンチマークは、チェックリストに基づくMLLM判定と専用の現実性評価基準を通じて、指示追従、忠実性保持、現実性、編集意図を評価し、さらに自動化されたクロスモーダル指標、ビデオ指標、音声指標によって補完される。広範な評価により、最先端モデルであっても非対象コンテンツを保持しながらクロスモーダル指示を実行することは依然として困難であることが示された。さらに我々は、複雑な指示を依存関係のあるサブタスクに分解し、自己内省と評価者フィードバックを通じて編集結果を反復的に改善するモジュール型エージェントフレームワークであるAVE-Agentを提案する。AVE-Agentは、競争力のある知覚品質を維持しつつ、統合編集における指示実行、忠実性保持、音声・視覚の整合性を向上させる。
English
While instruction-based video editing has advanced rapidly, real-world videos contain tightly coupled audio and visual signals, and editing one modality often requires coordinated changes in the other. Existing benchmarks primarily evaluate visual transformations on silent clips or isolated audio editing, leaving complex audio-visual editing and cross-modal consistency underexplored. We introduce AVE-Compass, a comprehensive benchmark with 145 curated source videos, 196 audio-visually coupled editing instructions, and 2,688 fine-grained checklist items. It evaluates Instruction Following, Fidelity Preserving, Realism, and Editing Intent through checklist-based MLLM judging and a dedicated realism rubric, complemented by automated cross-modal, video, and audio metrics. Extensive evaluation shows that state-of-the-art models still struggle to execute cross-modal instructions while preserving non-target content. We further propose AVE-Agent, a modular agent framework that decomposes complex instructions into dependent subtasks and iteratively improves editing results through self-reflection and evaluator feedback. AVE-Agent improves instruction execution, Fidelity Preserving, and audio-visual alignment in joint editing while maintaining competitive perceptual quality.