ChatPaper.aiChatPaper

ショットを考える:エージェント的推論による一貫したマルチショット動画編集

Thinking on Shots: Consistent Multi-Shot Video Editing with Agentic Reasoning

August 27, 2026
著者: Chenyang Wu, Fuchen Long, Binyuan Huang, Xinlong Sun, Xi Chen, Chun-Le Guo, Chongyi Li
cs.AI

要旨

生成AIはビデオ編集を大幅に進歩させたが、既存の手法は主にシングルショットまたは短いビデオクリップに焦点を当てている。複数の指示を含む長尺動画の編集は、依然として極めて困難な課題である。例えば固定長セグメンテーションのような単純なチャンキング戦略は、しばしばエンティティの断片化、深刻な編集ハルシネーション、時間的連続性の破壊を引き起こす。このギャップを埋めるために、我々はマルチ指示・マルチショット長尺ビデオ編集(MMLVE)タスクを提案する。これは、クロスショット編集一貫性(CSEC)、マルチ指示デカップリング(MID)、時空間構造に対するゼロ破壊(ZDSS)という3つの中核的な目的に基づいて構成されている。これらの3つの独自の課題に取り組むために、我々は大規模言語モデル(LLM)と視覚言語モデル(VLM)の相乗効果を活用し、ショットレベルのビデオ分離と正確な指示解析を実現するエージェント型編集フレームワークを導入する。さらに、このタスクを包括的に評価するために、複雑な実世界の時空間ダイナミクス、高密度で異種混合の指示、疎でランダムなエンティティ分布を特徴とするMMLVEに特化したデータセットであるMMLVE-Benchを構築する。編集結果の品質を評価するために、MMLVEに特化した3つの評価指標をさらに活用する。大規模な実験により、我々のMMLVE-Agentは既存のクローズドソースのSOTA手法(例:Seedance 2.0)を上回り、編集ハルシネーションの排除、クロスショット編集一貫性の維持、シームレスな時空間遷移の達成に成功することを実証する。
English
While generative AI has significantly advanced video editing, existing methods primarily focus on single-shot or short video clips. Editing long videos with multiple instructions remains a formidable challenge. Naive chunking strategies, e.g., fixed-duration segmentation, often lead to entity fragmentation, severe editing hallucinations, and disrupted temporal continuity. To bridge this gap, we introduce the Multi-Instruction Multi-Shot Long-Video Editing (MMLVE) task, which is structured around three core objectives: Cross-Shot Editing Consistency (CSEC), Multi-Instruction Decoupling (MID), and Zero-Destruction on Spatiotemporal Structure (ZDSS). To tackle these three unique challenges, we introduce an agentic editing framework that leverages the synergy of Large Language Models (LLMs) and Vision-Language Models (VLMs) to achieve shot-level video decoupling and precise instruction parsing. Furthermore, to comprehensively evaluate this task, we construct MMLVE-Bench, which is an MMLVE-focused dataset characterized by complex real-world spatiotemporal dynamics, high-density heterogeneous instructions, and sparse, random entity distributions. Three MMLVE-focused evaluation metrics are further exploited to assess the quality of the editing results. Extensive experiments demonstrate that our MMLVE-Agent outperforms existing closed-source SOTA approaches (e.g., Seedance 2.0), successfully eliminating editing hallucinations, preserving cross-shot editing consistency, and attaining seamless spatiotemporal transitions.