FlowMimic: マスク不要の視覚編集と生成 - ピクセルペアワープフローフィールドを用いたオンラインビデオ編集データ生成とモダリティ模倣
FlowMimic: Mask-free Visual Editing and Generation with Pixel-pair Warped Flow Field for Online Video Editing Data Generation and Modality Mimicry
July 20, 2026
著者: Dingyun Zhang, Lixue Gong, Wei Liu
cs.AI
要旨
視覚研究の主流方向に沿い、我々は動画と画像という二つのモダリティに対する生成機能と編集機能を単一モデル内に統合することを探求する。現在の動画編集データ収集手法は、物体マスクのアノテーション、I2VモデルとControlNet的なガイダンスによる誤差を含むペア合成の利用、VLMに基づく品質フィルタリングや改良といった、労力と時間を要する厳選された手順に依存しており、タスクの拡張性が限られている。その結果、編集タスクの多様性は画像編集モデルに比べて著しく狭いままである。我々は、画像編集サンプルから対応する動画編集サンプルをリアルタイムで直接生成できるピクセルペア時間ワープフローフィールドを開発し、複数レベルの動画編集タスクにわたって、モデルがこのようなデータのみを用いて動画編集を学習可能であることを実証する。我々は画像モダリティを動画モダリティの特殊な形態とみなす。これに基づき、モダリティ模倣生成損失とモダリティ模倣編集損失を設計し、相互模倣を通じて二つのモダリティの能力、ひいては出力分布を相対的に整合させる。さらに、言語ベースのビジュアル編集は、編集指示と参照ビジュアルコンテンツの理解、参照ビジュアルコンテンツ内でその指示に対応する領域の特定、そしてその領域のみの修正を伴う。既存の手法は主に外部の助けに依存しており、例えば追加のMLLMのファインチューニングや、推論時に補助入力としてマスクシーケンスを明示的に提供するといった方法がとられる。これに対し、我々はモデルがこの能力を内在化することを目指す。そのために、指示表現セグメンテーションなどの意味関連タスクとともに、対応する編集領域認識の潜在レベル損失とアテンションレベル損失を導入する。
English
In line with the prevailing direction of vision research, we explore the integration of both generation and editing capabilities for video and image modalities within a single model. Current approaches to collecting video editing data typically depend on labour-intensive, time-consuming curated procedures--involving object mask annotation, the use of error-introducing pair synthesis via I2V model and ControlNet-like guidance, and VLM-based quality filtering or refinement--and demonstrate limited task scalability. As a result, the diversity of editing tasks remains substantially narrower than that available for image editing models. We develop a pixel-pair temporal warped flow field that can directly generate corresponding video editing samples in real time from image editing samples, and we demonstrate across multiple levels of video editing tasks that a model can learn video editing using only such data. We regard the image modality as a particular form of the video modality. Accordingly, we design a modality mimic generation loss and a modality mimic editing loss to relatively align the capabilities--and thereby the output distributions--of the two modalities through mutual imitation. Moreover, language-based visual editing entails the comprehension of the editing instruction and the reference visual content, the localization of the region corresponding to that instruction within the reference visual contents, and the modification of that region alone. Existing approaches predominantly rely on external aids, such as fine-tuning an additional MLLM or explicitly supplying a mask sequence as auxiliary input during inference. In contrast, we aspire for the model to internalize this capability. To that end, we introduce sense-related tasks--for instance, referring expression segmentation--along with corresponding editing-region-aware latent-level loss and attention-level loss.