ChatPaper.aiChatPaper

FlowMimic:基于像素对扭曲流场的无掩模视觉编辑与生成方法,用于在线视频编辑数据生成与模态模仿

FlowMimic: Mask-free Visual Editing and Generation with Pixel-pair Warped Flow Field for Online Video Editing Data Generation and Modality Mimicry

July 20, 2026
作者: Dingyun Zhang, Lixue Gong, Wei Liu
cs.AI

摘要

按照视觉研究的主流方向,我们探索在单一模型中整合视频和图像模态的生成与编辑能力。当前收集视频编辑数据的方法通常依赖费时费力的人工标注流程,包括对象掩码标注、通过I2V模型和类似ControlNet的引导生成易引入错误的配对合成数据,以及基于VLM的质量过滤或精炼,且任务扩展性有限。因此,编辑任务的多样性远不如图像编辑模型。我们开发了一种像素对时间扭曲流场,能够直接从图像编辑样本实时生成对应的视频编辑样本,并在多个层次的视频编辑任务中证明,模型仅使用此类数据即可学习视频编辑。我们将图像模态视为视频模态的一种特殊形式。相应地,我们设计了模态模仿生成损失和模态模仿编辑损失,通过相互模仿来相对对齐两种模态的能力——进而对齐其输出分布。此外,基于语言的视觉编辑需要理解编辑指令和参考视觉内容,在参考视觉内容中定位指令对应的区域,并仅修改该区域。现有方法主要依赖外部辅助,例如微调额外的多模态大语言模型,或在推理时明确提供掩码序列作为辅助输入。相比之下,我们希望模型内化这种能力。为此,我们引入了与感知相关的任务——例如指代分割——以及相应的编辑区域感知的潜在级损失和注意力级损失。
English
In line with the prevailing direction of vision research, we explore the integration of both generation and editing capabilities for video and image modalities within a single model. Current approaches to collecting video editing data typically depend on labour-intensive, time-consuming curated procedures--involving object mask annotation, the use of error-introducing pair synthesis via I2V model and ControlNet-like guidance, and VLM-based quality filtering or refinement--and demonstrate limited task scalability. As a result, the diversity of editing tasks remains substantially narrower than that available for image editing models. We develop a pixel-pair temporal warped flow field that can directly generate corresponding video editing samples in real time from image editing samples, and we demonstrate across multiple levels of video editing tasks that a model can learn video editing using only such data. We regard the image modality as a particular form of the video modality. Accordingly, we design a modality mimic generation loss and a modality mimic editing loss to relatively align the capabilities--and thereby the output distributions--of the two modalities through mutual imitation. Moreover, language-based visual editing entails the comprehension of the editing instruction and the reference visual content, the localization of the region corresponding to that instruction within the reference visual contents, and the modification of that region alone. Existing approaches predominantly rely on external aids, such as fine-tuning an additional MLLM or explicitly supplying a mask sequence as auxiliary input during inference. In contrast, we aspire for the model to internalize this capability. To that end, we introduce sense-related tasks--for instance, referring expression segmentation--along with corresponding editing-region-aware latent-level loss and attention-level loss.