ChatPaper.aiChatPaper

CoinVE-200K: 構成的指示誘導型ビデオ編集のための大規模高品質データセット

CoinVE-200K: A Large-Scale High-Quality Dataset for Compositional Instruction-Guided Video Editing

August 18, 2026
著者: Fuchen Long, Cong Wang, Zitao Gao, Wenhao Zhong, Yu Cheng, Xiaolu Hou, Yan Li, Xiao Cao, Xinlong Sun, Xi Chen, Yu Liu
cs.AI

要旨

指示ベースのビデオ編集データセットの品質と多様性は着実に向上しているが、既存のデータセットは主に単一の編集操作に焦点を当てており、構成的指示誘導ビデオ編集をサポートするには不十分である。特に、複数の編集意図を同じビデオ内で同時に理解し、忠実に実行する必要がある。この問題に対処するため、我々は構成的指示誘導ビデオ編集のための大規模かつ高品質なデータセットであるCoinVE-200Kを導入する。CoinVE-200Kには、最大201フレームの1080pビデオ編集ペアが含まれており、各サンプルが2〜5個の原子的編集操作を含む多様な構成的シナリオを網羅している。指示は人間、物体、背景を対象とし、追加、削除、修正、スタイライゼーションなどの編集タイプをカバーしている。すべてのサンプルは、指示忠実性、視覚的品質、時間的一貫性、構成的多様性を保証するために、慎重に設計された生成・フィルタリングパイプラインを通じて構築されている。また、多様な対象、操作タイプ、指示の複雑さにわたる構成的指示ビデオ編集のためのベンチマークであるCoinVE-Benchも導入する。さらに、Wan2.1-T2V-14BとQwen3-VL-8B-Instructに基づいて構築された22Bの構成的ビデオ編集モデルであるCoinVE-Editを提示する。CoinVE-Editは、異なる編集指示に対して領域認識アテンションを分離し、無関係なコンテンツと時間的整合性を保ちながら、正確な複数領域編集を可能にする。CoinVE-Benchでの実験は、CoinVE-Editが指示追従、構成的編集精度、視覚的品質、時間的一貫性において高い性能を達成することを示している。
English
The quality and diversity of instruction-based video editing datasets are steadily improving, yet existing datasets mainly focus on single editing operations and fall short in supporting compositional instruction-guided video editing. In particular, multiple editing intents must be jointly understood and faithfully executed within the same video. To address this issue, we introduce CoinVE-200K, a large-scale, high-quality dataset for Compositional Instruction-Guided Video Editing. CoinVE-200K contains 1080p video-editing pairs of up to 201 frames, covering diverse compositional scenarios where each sample involves 2 to 5 atomic editing operations. The instructions target humans, objects, and backgrounds, and cover edit types such as addition, removal, modification, and stylization. All samples are built through a carefully designed generation and filtering pipeline to ensure instruction faithfulness, visual quality, temporal consistency, and compositional diversity. We also introduce CoinVE-Bench, a benchmark for compositional-instruction video editing across diverse subjects, operation types, and instruction complexities. Furthermore, we present CoinVE-Edit, a 22B compositional video editing model built upon Wan2.1-T2V-14B and Qwen3-VL-8B-Instruct. CoinVE-Edit disentangles region-aware attention for different editing instructions, enabling precise multi-region editing while preserving irrelevant content and temporal coherence. Experiments on CoinVE-Bench show that CoinVE-Edit achieves strong performance in instruction following, compositional editing accuracy, visual quality, and temporal consistency.