ChatPaper.aiChatPaper

FlowMimic:以像素對扭曲流場實現免遮罩視覺編輯與生成,應用於線上影片編輯資料生成與模態模仿

FlowMimic: Mask-free Visual Editing and Generation with Pixel-pair Warped Flow Field for Online Video Editing Data Generation and Modality Mimicry

July 20, 2026
作者: Dingyun Zhang, Lixue Gong, Wei Liu
cs.AI

摘要

遵循视觉研究的主流方向,我們探索在單一模型中整合影片與影像模態的生成與編輯能力。現有收集影片編輯資料的方法通常依賴耗時費力的人工整理流程——包括物件遮罩標註、透過 I2V 模型與 ControlNet 類引導進行易出錯的配對合成,以及基於 VLM 的品質過濾或精煉——且展現出有限的可擴展性。因此,影片編輯任務的多樣性仍遠窄於影像編輯模型可達到的範圍。我們開發了一種基於像素對應的時序扭曲流場,可從影像編輯樣本即時生成對應的影片編輯樣本,並在多個層次的影片編輯任務中證明,模型僅需此類資料即可學習影片編輯。我們將影像模態視為影片模態的一種特殊形式,因此設計了模態模仿生成損失與模態模仿編輯損失,透過相互模仿來相對對齊兩種模態的能力(進而對齊其輸出分布)。此外,基於語言的視覺編輯需理解編輯指令與參考視覺內容、定位指令對應的區域,並僅修改該區域。現有方法多依賴外部輔助,例如微調額外的 MLLM 或在推理時明確提供遮罩序列作為輔助輸入。相比之下,我們期望模型能內化此能力。為此,我們引入了感知相關任務(如指代表達分割)與對應的編輯區域感知潛在層損失及注意力層損失。
English
In line with the prevailing direction of vision research, we explore the integration of both generation and editing capabilities for video and image modalities within a single model. Current approaches to collecting video editing data typically depend on labour-intensive, time-consuming curated procedures--involving object mask annotation, the use of error-introducing pair synthesis via I2V model and ControlNet-like guidance, and VLM-based quality filtering or refinement--and demonstrate limited task scalability. As a result, the diversity of editing tasks remains substantially narrower than that available for image editing models. We develop a pixel-pair temporal warped flow field that can directly generate corresponding video editing samples in real time from image editing samples, and we demonstrate across multiple levels of video editing tasks that a model can learn video editing using only such data. We regard the image modality as a particular form of the video modality. Accordingly, we design a modality mimic generation loss and a modality mimic editing loss to relatively align the capabilities--and thereby the output distributions--of the two modalities through mutual imitation. Moreover, language-based visual editing entails the comprehension of the editing instruction and the reference visual content, the localization of the region corresponding to that instruction within the reference visual contents, and the modification of that region alone. Existing approaches predominantly rely on external aids, such as fine-tuning an additional MLLM or explicitly supplying a mask sequence as auxiliary input during inference. In contrast, we aspire for the model to internalize this capability. To that end, we introduce sense-related tasks--for instance, referring expression segmentation--along with corresponding editing-region-aware latent-level loss and attention-level loss.