AVE-Compass: 오디오-비디오 편집 능력에 대한 종합적 평가를 향하여
AVE-Compass: Towards Holistic Evaluation for Audio-Video Editing Abilities
July 17, 2026
저자: Yuqing Wen, Yukai Huang, Qianqian Xie, Jiangtao Wu, Yibin Lin, Yikai Gu, Jialu Chen, Yuanxing Zhang, Jiaheng Liu
cs.AI
초록
명령 기반 비디오 편집이 빠르게 발전해 왔지만, 실제 세계의 비디오는 오디오와 시각 신호가 긴밀하게 결합되어 있어 한 모달리티를 편집할 때 다른 모달리티의 조정이 함께 필요한 경우가 많다. 기존 벤치마크는 주로 무음 클립에 대한 시각적 변환이나 단독 오디오 편집을 평가하여, 복잡한 오디오-비주얼 편집과 교차 모달리티 일관성은 충분히 탐구되지 않았다. 본 연구에서는 145개의 엄선된 소스 비디오, 196개의 오디오-비주얼 결합 편집 명령, 2,688개의 세분화된 체크리스트 항목으로 구성된 종합 벤치마크 AVE-Compass를 소개한다. AVE-Compass는 체크리스트 기반 MLLM 평가와 전용 현실감 루브릭을 통해 명령 수행, 충실도 유지, 현실감, 편집 의도를 평가하며, 자동화된 교차 모달리티, 비디오, 오디오 메트릭으로 보완된다. 광범위한 평가 결과, 최첨단 모델들은 비대상 콘텐츠를 보존하면서 교차 모달리티 명령을 실행하는 데 여전히 어려움을 겪는 것으로 나타났다. 또한, 복잡한 명령을 종속적 하위 작업으로 분해하고 자기 반성과 평가자 피드백을 통해 편집 결과를 반복적으로 개선하는 모듈형 에이전트 프레임워크인 AVE-Agent를 제안한다. AVE-Agent는 공동 편집에서 명령 실행, 충실도 유지, 오디오-비주얼 정렬을 개선하면서 경쟁력 있는 지각 품질을 유지한다.
English
While instruction-based video editing has advanced rapidly, real-world videos contain tightly coupled audio and visual signals, and editing one modality often requires coordinated changes in the other. Existing benchmarks primarily evaluate visual transformations on silent clips or isolated audio editing, leaving complex audio-visual editing and cross-modal consistency underexplored. We introduce AVE-Compass, a comprehensive benchmark with 145 curated source videos, 196 audio-visually coupled editing instructions, and 2,688 fine-grained checklist items. It evaluates Instruction Following, Fidelity Preserving, Realism, and Editing Intent through checklist-based MLLM judging and a dedicated realism rubric, complemented by automated cross-modal, video, and audio metrics. Extensive evaluation shows that state-of-the-art models still struggle to execute cross-modal instructions while preserving non-target content. We further propose AVE-Agent, a modular agent framework that decomposes complex instructions into dependent subtasks and iteratively improves editing results through self-reflection and evaluator feedback. AVE-Agent improves instruction execution, Fidelity Preserving, and audio-visual alignment in joint editing while maintaining competitive perceptual quality.