ChatPaper.aiChatPaper

FlowMimic: 온라인 비디오 편집 데이터 생성 및 모달리티 모방을 위한 픽셀 쌍 왜곡 플로우 필드 기반의 마스크 없는 시각적 편집 및 생성

FlowMimic: Mask-free Visual Editing and Generation with Pixel-pair Warped Flow Field for Online Video Editing Data Generation and Modality Mimicry

July 20, 2026
저자: Dingyun Zhang, Lixue Gong, Wei Liu
cs.AI

초록

시각 연구의 주류 방향에 맞춰, 우리는 비디오와 이미지 모달리티에 대한 생성 및 편집 기능을 단일 모델 내에서 통합하는 방안을 탐구한다. 현재 비디오 편집 데이터 수집 접근 방식은 일반적으로 노동 집약적이고 시간이 많이 소요되는 큐레이션 절차(객체 마스크 주석, I2V 모델 및 ControlNet 유사 가이던스를 통한 오류 유발 쌍 합성, VLM 기반 품질 필터링 또는 개선 등)에 의존하며, 작업 확장성이 제한적이다. 그 결과, 편집 작업의 다양성은 이미지 편집 모델에서 사용 가능한 것보다 훨씬 좁은 수준에 머물러 있다. 우리는 이미지 편집 샘플로부터 실시간으로 해당 비디오 편집 샘플을 직접 생성할 수 있는 픽셀 쌍 시간적 워프 플로우 필드를 개발하며, 다양한 수준의 비디오 편집 작업에 걸쳐 모델이 이러한 데이터만으로도 비디오 편집을 학습할 수 있음을 입증한다. 우리는 이미지 모달리티를 비디오 모달리티의 특수한 형태로 간주한다. 이에 따라, 우리는 모달리티 모방 생성 손실과 모달리티 모방 편집 손실을 설계하여 상호 모방을 통해 두 모달리티의 능력, 나아가 출력 분포를 상대적으로 정렬한다. 또한, 언어 기반 시각 편집은 편집 명령과 참조 시각 콘텐츠의 이해, 참조 시각 콘텐츠 내에서 해당 명령에 대응하는 영역의 위치 파악, 그리고 그 영역만의 수정을 수반한다. 기존 접근 방식은 주로 추가 MLLM을 미세 조정하거나 추론 중에 보조 입력으로 마스크 시퀀스를 명시적으로 제공하는 등의 외부 도구에 의존한다. 반면, 우리는 모델이 이 능력을 내재화하기를 지향한다. 이를 위해, 우리는 의미 관련 작업(예: 지시적 표현 분할)과 함께 대응하는 편집 영역 인식 잠재 수준 손실 및 어텐션 수준 손실을 도입한다.
English
In line with the prevailing direction of vision research, we explore the integration of both generation and editing capabilities for video and image modalities within a single model. Current approaches to collecting video editing data typically depend on labour-intensive, time-consuming curated procedures--involving object mask annotation, the use of error-introducing pair synthesis via I2V model and ControlNet-like guidance, and VLM-based quality filtering or refinement--and demonstrate limited task scalability. As a result, the diversity of editing tasks remains substantially narrower than that available for image editing models. We develop a pixel-pair temporal warped flow field that can directly generate corresponding video editing samples in real time from image editing samples, and we demonstrate across multiple levels of video editing tasks that a model can learn video editing using only such data. We regard the image modality as a particular form of the video modality. Accordingly, we design a modality mimic generation loss and a modality mimic editing loss to relatively align the capabilities--and thereby the output distributions--of the two modalities through mutual imitation. Moreover, language-based visual editing entails the comprehension of the editing instruction and the reference visual content, the localization of the region corresponding to that instruction within the reference visual contents, and the modification of that region alone. Existing approaches predominantly rely on external aids, such as fine-tuning an additional MLLM or explicitly supplying a mask sequence as auxiliary input during inference. In contrast, we aspire for the model to internalize this capability. To that end, we introduce sense-related tasks--for instance, referring expression segmentation--along with corresponding editing-region-aware latent-level loss and attention-level loss.