鏡頭思考:具備代理式推理的一致性多鏡頭影片編輯
Thinking on Shots: Consistent Multi-Shot Video Editing with Agentic Reasoning
August 27, 2026
作者: Chenyang Wu, Fuchen Long, Binyuan Huang, Xinlong Sun, Xi Chen, Chun-Le Guo, Chongyi Li
cs.AI
摘要
儘管生成式人工智慧已顯著推動影片編輯的發展,現有方法仍主要專注於單鏡頭或短影片片段。利用多條指令編輯長影片依然是一項嚴峻的挑戰。樸素的分塊策略(如固定時長分割)往往導致實體碎片化、嚴重的編輯幻覺以及時間連續性中斷。為填補此一空白,我們提出了多指令多鏡頭長影片編輯(Multi-Instruction Multi-Shot Long-Video Editing, MMLVE)任務,該任務圍繞三個核心目標建構:跨鏡頭編輯一致性(Cross-Shot Editing Consistency, CSEC)、多指令解耦(Multi-Instruction Decoupling, MID)與時空結構零破壞(Zero-Destruction on Spatiotemporal Structure, ZDSS)。為應對這三項獨特挑戰,我們提出了一種智慧體編輯框架,利用大型語言模型(Large Language Models, LLMs)與視覺語言模型(Vision-Language Models, VLMs)之協同作用,實現鏡頭層級的影片解耦及精確的指令解析。此外,為全面評估此任務,我們建構了MMLVE-Bench——一個以MMLVE為核心的資料集,其特徵在於複雜的真實世界時空動態、高密度異質指令,以及稀疏且隨機的實體分佈。我們並進一步採用三項針對MMLVE的評估指標,以衡量編輯結果的品質。大量實驗證實,我們的MMLVE-Agent優於現有的閉源最先進(SOTA)方法(如Seedance 2.0),成功消除了編輯幻覺、保持了跨鏡頭編輯一致性,並實現了無縫的時空過渡。
English
While generative AI has significantly advanced video editing, existing methods primarily focus on single-shot or short video clips. Editing long videos with multiple instructions remains a formidable challenge. Naive chunking strategies, e.g., fixed-duration segmentation, often lead to entity fragmentation, severe editing hallucinations, and disrupted temporal continuity. To bridge this gap, we introduce the Multi-Instruction Multi-Shot Long-Video Editing (MMLVE) task, which is structured around three core objectives: Cross-Shot Editing Consistency (CSEC), Multi-Instruction Decoupling (MID), and Zero-Destruction on Spatiotemporal Structure (ZDSS). To tackle these three unique challenges, we introduce an agentic editing framework that leverages the synergy of Large Language Models (LLMs) and Vision-Language Models (VLMs) to achieve shot-level video decoupling and precise instruction parsing. Furthermore, to comprehensively evaluate this task, we construct MMLVE-Bench, which is an MMLVE-focused dataset characterized by complex real-world spatiotemporal dynamics, high-density heterogeneous instructions, and sparse, random entity distributions. Three MMLVE-focused evaluation metrics are further exploited to assess the quality of the editing results. Extensive experiments demonstrate that our MMLVE-Agent outperforms existing closed-source SOTA approaches (e.g., Seedance 2.0), successfully eliminating editing hallucinations, preserving cross-shot editing consistency, and attaining seamless spatiotemporal transitions.