video-generation image-editing video-editing pixel-pair latent-loss attention-loss referring-expression editing-instruction
Abstract
In line with the prevailing direction of vision research, we explore the integration of both generation and editing capabilities for video and image modalities within a single model. Current approaches to collecting video editing data typically depend on labour-intensive, time-consuming curated procedures--involving object mask annotation, the use of error-introducing pair synthesis via I2V model and ControlNet-like guidance, and VLM-based quality filtering or refinement--and demonstrate limited task scalability. As a result, the diversity of editing tasks remains substantially narrower than that available for image editing models. We develop a pixel-pair temporal warped flow field that can directly generate corresponding video editing samples in real time from image editing samples, and we demonstrate across multiple levels of video editing tasks that a model can learn video editing using only such data. We regard the image modality as a particular form of the video modality. Accordingly, we design a modality mimic generation loss and a modality mimic editing loss to relatively align the capabilities--and thereby the output distributions--of the two modalities through mutual imitation. Moreover, language-based visual editing entails the comprehension of the editing instruction and the reference visual content, the localization of the region corresponding to that instruction within the reference visual contents, and the modification of that region alone. Existing approaches predominantly rely on external aids, such as fine-tuning an additional MLLM or explicitly supplying a mask sequence as auxiliary input during inference. In contrast, we aspire for the model to internalize this capability. To that end, we introduce sense-related tasks--for instance, referring expression segmentation--along with corresponding editing-region-aware latent-level loss and attention-level loss.
한국어 요약
한 줄 요약
FlowMimic은 마스크 없이 영상 편집 및 생성을 가능하게 하는 픽셀-쌍 왜곡 흐름 필드를 제안하여 영상 편집 데이터 생성과 모달리티 모방을 실현한다.
핵심 기여도
- 픽셀-쌍 기반 템포럴 왜곡 흐름 필드를 도입하여 이미지 편집 샘플에서 실시간 영상 편집 샘플 생성 가능.
- 모달리티 모방 생성 손실(modality mimic generation loss)과 편집 손실(modality mimic editing loss)을 통해 이미지와 영상 모델의 출력 분포를 상호 정렬.
- 외부 도구 없이 편집 영역을 자동으로 인식하는 레퍼런스 표현 분할과 레이턴트-레벨, 어텐션-레벨 손실을 제안.
핵심 아이디어
FlowMimic은 영상 편집 데이터 생성 과정에서 수작업 마스크 또는 외부 모델에 의존하는 기존 방식의 한계를 극복하고자 한다. 기존 연구는 편집 영역을 명시하기 위해 마스크를 수작업하거나, I2V 모델과 ControlNet과 같은 외부 가이드라인을 사용하는 방식으로, 시간과 노동이 많이 드는 문제가 있었다. FlowMimic은 대신 픽셀-쌍을 기반으로 한 템포럴 왜곡 흐름 필드를 도입하여, 이미지 편집 샘플을 실시간으로 영상 편집 샘플로 변환할 수 있도록 한다. 또한, 이미지를 영상의 특별한 형태로 간주하고, 두 모달리티 간의 분포 정렬을 위해 모달리티 모방 손실을 설계하여 모델의 내재적 편집 능력을 강화한다.
기술적 접근법
- 픽셀-쌍 기반 템포럴 왜곡 흐름 필드를 사용하여 이미지 편집 샘플에서 영상 편집 샘플을 생성.
- 모달리티 모방 생성 손실과 편집 손실을 통해 이미지와 영상 모델의 출력 분포를 정렬.
- 레퍼런스 표현 분할과 같은 감각 관련 작업을 통해 편집 영역을 자동 인식.
- 레이턴트-레벨 손실과 어텐션-레벨 손실을 통해 편집 영역에 대한 모델의 이해를 강화.
주요 결과
- 다양한 영상 편집 작업에서 FlowMimic 모델이 단일 이미지 편집 샘플만으로 학습 가능함을 실험적으로 입증.
- 외부 마스크 또는 추가 MLLM 없이도 편집 영역을 정확히 인식하는 능력을 보임.
- 기존 방법 대비 편집 데이터 생성 과정에서 수작업 및 외부 도구 사용을 최소화함으로써 작업 효율성 향상.
의의 및 한계
FlowMimic은 영상 편집 데이터 생성 과정의 자동화와 모델 내재적 편집 능력 강화라는 두 가지 측면에서 중요한 기여를 한다. 특히, 이미지 편집 샘플만으로 영상 편집 모델을 학습할 수 있다는 점에서 기존 연구와 차별화된다. 그러나 특정 복잡한 편집 작업에서는 외부 마스크나 가이드라인 없이도 충분한 성능을 보장하기 어려울 수 있다. 또한, 다양한 영상 편집 작업에서의 일반화 가능성에 대한 추가 실험 필요성이 제기된다.
실용적 활용
FlowMimic은 온라인 영상 편집 플랫폼, 자동 편집 도구, 영상 생성 AI 등에서 활용될 수 있으며, 특히 편집 데이터 생성 과정의 자동화를 통해 제작 비용을 절감할 수 있다. 또한, 편집 영역 인식 기능은 영상 콘텐츠 제작, 광고, 콘텐츠 편집 등 다양한 산업 분야에서 유용하게 활용될 수 있다.