샷에 대한 추론: 에이전트 기반 추론을 통한 일관된 멀티샷 비디오 편집
Thinking on Shots: Consistent Multi-Shot Video Editing with Agentic Reasoning
August 27, 2026
저자: Chenyang Wu, Fuchen Long, Binyuan Huang, Xinlong Sun, Xi Chen, Chun-Le Guo, Chongyi Li
cs.AI
초록
생성형 AI가 비디오 편집 분야를 크게 발전시켰지만, 기존 방법은 주로 단일 샷 또는 짧은 비디오 클립에 초점을 맞추고 있다. 여러 지시사항이 포함된 장편 비디오 편집은 여전히 난제로 남아 있다. 고정 길이 분할과 같은 단순한 청크 분할 전략은 종종 개체 단편화, 심각한 편집 환각, 시간적 연속성 붕괴를 초래한다. 이러한 격차를 해소하기 위해, 우리는 세 가지 핵심 목표를 중심으로 구조화된 다중 지시 다중 샷 장편 비디오 편집(MMLVE) 작업을 소개한다: 크로스 샷 편집 일관성(CSEC), 다중 지시 디커플링(MID), 시공간 구조 무손상(ZDSS). 이러한 세 가지 고유한 과제를 해결하기 위해, 우리는 대규모 언어 모델(LLM)과 비전-언어 모델(VLM)의 시너지를 활용하여 샷 수준 비디오 디커플링과 정밀한 지시 파싱을 달성하는 에이전트 기반 편집 프레임워크를 도입한다. 또한, 이 작업을 종합적으로 평가하기 위해 복잡한 실제 시공간 역학, 고밀도 이질적 지시, 희소하고 무작위적인 개체 분포를 특징으로 하는 MMLVE 중심 데이터셋인 MMLVE-Bench를 구축한다. 편집 결과의 품질을 평가하기 위해 세 가지 MMLVE 중심 평가 지표를 추가로 활용한다. 광범위한 실험을 통해 우리의 MMLVE-Agent가 Seedance 2.0과 같은 기존 폐쇄 소스 최고 수준(SOTA) 접근 방식을 능가하며, 편집 환각을 성공적으로 제거하고 크로스 샷 편집 일관성을 유지하며 매끄러운 시공간 전환을 달성함을 입증한다.
English
While generative AI has significantly advanced video editing, existing methods primarily focus on single-shot or short video clips. Editing long videos with multiple instructions remains a formidable challenge. Naive chunking strategies, e.g., fixed-duration segmentation, often lead to entity fragmentation, severe editing hallucinations, and disrupted temporal continuity. To bridge this gap, we introduce the Multi-Instruction Multi-Shot Long-Video Editing (MMLVE) task, which is structured around three core objectives: Cross-Shot Editing Consistency (CSEC), Multi-Instruction Decoupling (MID), and Zero-Destruction on Spatiotemporal Structure (ZDSS). To tackle these three unique challenges, we introduce an agentic editing framework that leverages the synergy of Large Language Models (LLMs) and Vision-Language Models (VLMs) to achieve shot-level video decoupling and precise instruction parsing. Furthermore, to comprehensively evaluate this task, we construct MMLVE-Bench, which is an MMLVE-focused dataset characterized by complex real-world spatiotemporal dynamics, high-density heterogeneous instructions, and sparse, random entity distributions. Three MMLVE-focused evaluation metrics are further exploited to assess the quality of the editing results. Extensive experiments demonstrate that our MMLVE-Agent outperforms existing closed-source SOTA approaches (e.g., Seedance 2.0), successfully eliminating editing hallucinations, preserving cross-shot editing consistency, and attaining seamless spatiotemporal transitions.