AVE-Compass:邁向音視頻編輯能力的全面評估
AVE-Compass: Towards Holistic Evaluation for Audio-Video Editing Abilities
July 17, 2026
作者: Yuqing Wen, Yukai Huang, Qianqian Xie, Jiangtao Wu, Yibin Lin, Yikai Gu, Jialu Chen, Yuanxing Zhang, Jiaheng Liu
cs.AI
摘要
儘管基於指令的影片編輯技術進展迅速,但真實世界影片中的音訊與視覺訊號緊密耦合,編輯其中一個模態往往需要對另一個模態進行協調變更。現有基準主要針對無聲片段上的視覺轉換或孤立的音訊編輯進行評估,使得複雜的視聽編輯與跨模態一致性探討不足。我們推出 AVE-Compass,一個全面的基準,包含 145 個精選來源影片、196 條視聽耦合的編輯指令,以及 2,688 個細粒度檢查表項目。它透過基於檢查表的多模態大語言模型評判與專門的真實感評分標準,評估指令跟隨、保真度維持、真實感與編輯意圖,並輔以自動化的跨模態、影片與音訊指標。廣泛的評估顯示,現有最佳模型在執行跨模態指令的同時仍難以保留非目標內容。我們進一步提出 AVE-Agent,一個模組化智慧體框架,能將複雜指令分解為相互依賴的子任務,並透過自我反思與評估者反饋迭代地改善編輯結果。AVE-Agent 在聯合編輯中提升了指令執行、保真度維持與視聽對齊,同時保持具有競爭力的感知品質。
English
While instruction-based video editing has advanced rapidly, real-world videos contain tightly coupled audio and visual signals, and editing one modality often requires coordinated changes in the other. Existing benchmarks primarily evaluate visual transformations on silent clips or isolated audio editing, leaving complex audio-visual editing and cross-modal consistency underexplored. We introduce AVE-Compass, a comprehensive benchmark with 145 curated source videos, 196 audio-visually coupled editing instructions, and 2,688 fine-grained checklist items. It evaluates Instruction Following, Fidelity Preserving, Realism, and Editing Intent through checklist-based MLLM judging and a dedicated realism rubric, complemented by automated cross-modal, video, and audio metrics. Extensive evaluation shows that state-of-the-art models still struggle to execute cross-modal instructions while preserving non-target content. We further propose AVE-Agent, a modular agent framework that decomposes complex instructions into dependent subtasks and iteratively improves editing results through self-reflection and evaluator feedback. AVE-Agent improves instruction execution, Fidelity Preserving, and audio-visual alignment in joint editing while maintaining competitive perceptual quality.