AVE-Compass:面向音视频编辑能力的全面评估
AVE-Compass: Towards Holistic Evaluation for Audio-Video Editing Abilities
July 17, 2026
作者: Yuqing Wen, Yukai Huang, Qianqian Xie, Jiangtao Wu, Yibin Lin, Yikai Gu, Jialu Chen, Yuanxing Zhang, Jiaheng Liu
cs.AI
摘要
尽管基于指令的视频编辑技术已取得快速发展,但真实世界视频中的音频与视觉信号紧密耦合,编辑其中一种模态往往需要另一种模态进行协同变化。现有基准主要评估静音片段上的视觉变换或孤立的音频编辑,导致复杂的音视频联合编辑与跨模态一致性仍缺乏充分探索。为此,我们提出 AVE-Compass,一个综合性基准,包含 145 个精选源视频、196 条音视频耦合编辑指令以及 2,688 个细粒度检查项。该基准通过基于检查项的多模态大语言模型评判和专门设计的真实感评分准则,评估指令跟随、保真度保持、真实感与编辑意图,并辅以自动化的跨模态、视频和音频指标。大量评估表明,现有最先进模型在执行跨模态指令的同时仍难以保持非目标内容不变。我们进一步提出 AVE-Agent,一个模块化智能体框架,可将复杂指令分解为相互依赖的子任务,并通过自我反思与评估器反馈迭代改进编辑结果。AVE-Agent 在联合编辑中显著提升了指令执行、保真度保持和音视频对齐能力,同时保持了具有竞争力的感知质量。
English
While instruction-based video editing has advanced rapidly, real-world videos contain tightly coupled audio and visual signals, and editing one modality often requires coordinated changes in the other. Existing benchmarks primarily evaluate visual transformations on silent clips or isolated audio editing, leaving complex audio-visual editing and cross-modal consistency underexplored. We introduce AVE-Compass, a comprehensive benchmark with 145 curated source videos, 196 audio-visually coupled editing instructions, and 2,688 fine-grained checklist items. It evaluates Instruction Following, Fidelity Preserving, Realism, and Editing Intent through checklist-based MLLM judging and a dedicated realism rubric, complemented by automated cross-modal, video, and audio metrics. Extensive evaluation shows that state-of-the-art models still struggle to execute cross-modal instructions while preserving non-target content. We further propose AVE-Agent, a modular agent framework that decomposes complex instructions into dependent subtasks and iteratively improves editing results through self-reflection and evaluator feedback. AVE-Agent improves instruction execution, Fidelity Preserving, and audio-visual alignment in joint editing while maintaining competitive perceptual quality.