基于镜头的思考:具有智能体推理的一致性多镜头视频编辑
Thinking on Shots: Consistent Multi-Shot Video Editing with Agentic Reasoning
August 27, 2026
作者: Chenyang Wu, Fuchen Long, Binyuan Huang, Xinlong Sun, Xi Chen, Chun-Le Guo, Chongyi Li
cs.AI
摘要
尽管生成式AI显著推动了视频编辑的发展,但现有方法主要聚焦于单镜头或短视频片段。使用多指令编辑长视频仍然是一项严峻的挑战。朴素的切分策略(例如固定时长分割)往往导致实体碎片化、严重的编辑幻觉以及时间连续性的破坏。为弥合这一差距,我们提出了多指令多镜头长视频编辑(MMLVE)任务,该任务围绕三个核心目标构建:跨镜头编辑一致性(CSEC)、多指令解耦(MID)和时空结构零破坏(ZDSS)。为应对这三个独特挑战,我们引入了一种智能体编辑框架,利用大语言模型(LLMs)与视觉语言模型(VLMs)的协同作用实现镜头级视频解耦和精确指令解析。此外,为全面评估该任务,我们构建了MMLVE-Bench——一个聚焦于MMLVE的数据集,其特点是包含复杂的真实世界时空动态、高密度异构指令以及稀疏随机的实体分布。我们进一步利用三个聚焦于MMLVE的评估指标来衡量编辑结果的质量。大量实验表明,我们的MMLVE-Agent优于现有的闭源SOTA方法(如Seedance 2.0),成功消除了编辑幻觉,保持了跨镜头编辑一致性,并实现了无缝的时空过渡。
English
While generative AI has significantly advanced video editing, existing methods primarily focus on single-shot or short video clips. Editing long videos with multiple instructions remains a formidable challenge. Naive chunking strategies, e.g., fixed-duration segmentation, often lead to entity fragmentation, severe editing hallucinations, and disrupted temporal continuity. To bridge this gap, we introduce the Multi-Instruction Multi-Shot Long-Video Editing (MMLVE) task, which is structured around three core objectives: Cross-Shot Editing Consistency (CSEC), Multi-Instruction Decoupling (MID), and Zero-Destruction on Spatiotemporal Structure (ZDSS). To tackle these three unique challenges, we introduce an agentic editing framework that leverages the synergy of Large Language Models (LLMs) and Vision-Language Models (VLMs) to achieve shot-level video decoupling and precise instruction parsing. Furthermore, to comprehensively evaluate this task, we construct MMLVE-Bench, which is an MMLVE-focused dataset characterized by complex real-world spatiotemporal dynamics, high-density heterogeneous instructions, and sparse, random entity distributions. Three MMLVE-focused evaluation metrics are further exploited to assess the quality of the editing results. Extensive experiments demonstrate that our MMLVE-Agent outperforms existing closed-source SOTA approaches (e.g., Seedance 2.0), successfully eliminating editing hallucinations, preserving cross-shot editing consistency, and attaining seamless spatiotemporal transitions.