vision-language flow-matching diffusion-transformer rectified-flow conditional-generation tool-based-editing image-retouching mmart-bench
Abstract
Tool-based image editing (image retouching) is commonly formulated with autoregressive multimodal large language models (MLLMs) that sequentially generate reasoning, tool selections, and parameter values. In this work, we present a novel approach to tool-based image editing by framing the task as a flow matching problem. We introduce FlowTool, a framework that directly models the distribution of high-quality tool parameters conditioned on the input image and user instruction using conditional rectified flow. FlowTool combines a vision-language model backbone for multimodal understanding with a Diffusion Transformer parameter generator that transforms Gaussian noise into an editing plan. We train FlowTool with a two-stage supervised flow-matching curriculum, followed by reward-based post-training. Across MMArt-Bench, FlowTool-Eval, ArtEdit-Bench, and MIT-Adobe5K, FlowTool achieves significantly stronger reference-based performance than specialized MLLM editing agents and proprietary MLLMs, while remaining competitive with proprietary models under reference-free evaluation. Moreover, FlowTool significantly improves inference efficiency, reducing latency by at least 50times while requiring nearly 2times less memory than the compared baselines. These results demonstrate that tool-based image editing can be effectively modeled as conditional generation over structured continuous editing parameters, without autoregressive reasoning.
한국어 요약
한 줄 요약
FlowTool은 이미지 리터치 작업에서 툴 파라미터를 조절하기 위해 흐름 매칭을 활용한 새로운 프레임워크이다.
핵심 기여도
- FlowTool은 조건부 정규화 흐름(conditional rectified flow)을 사용하여 입력 이미지와 사용자 지시에 조건을 주고, 고품질 툴 파라미터 분포를 직접 모델링한다.
- Diffusion Transformer 파라미터 생성기를 도입하여 가우시안 노이즈를 편집 계획으로 변환한다.
- 투자적 흐름 매칭 커리큘럼과 보상 기반 사후 훈련을 통해 FlowTool을 학습한다.
- MMArt-Bench, FlowTool-Eval, ArtEdit-Bench, MIT-Adobe5K 데이터셋에서 기존 MLLM 에이전트 대비 우수한 성능을 보인다.
핵심 아이디어
기존 이미지 리터치는 자동 회귀적 MLLM을 사용해 툴 선택과 파라미터 생성을 순차적으로 처리했으나, 이는 추론 단계가 복잡하고 비효율적이다. FlowTool은 이러한 문제를 해결하기 위해 툴 파라미터 생성을 흐름 매칭 문제로 재구성한다. 이 접근법은 조건부 정규화 흐름을 통해 이미지와 사용자 지시를 기반으로 연속적인 편집 파라미터를 생성함으로써, 자동 회귀적 추론 없이도 효과적인 편집 계획을 도출할 수 있다. Diffusion Transformer는 노이즈를 편집 계획으로 변환하는 역할을 하며, 이는 기존 방식보다 구조화된 파라미터 공간을 직접 생성할 수 있다는 장점을 가진다.
기술적 접근법
- **모델 구조**: Vision-Language 모델(VLM) 백본과 Diffusion Transformer 파라미터 생성기를 결합.
- **학습 전략**: 두 단계의 감독 흐름 매칭 커리큘럼 + 보상 기반 사후 훈련.
- **입력 조건**: 입력 이미지와 사용자 지시에 조건을 주어 파라미터 분포를 모델링.
- **파라미터 생성**: 가우시안 노이즈를 편집 계획으로 변환하는 Diffusion Transformer 사용.
주요 결과
- MMArt-Bench, FlowTool-Eval, ArtEdit-Bench, MIT-Adobe5K 데이터셋에서 기존 MLLM 에이전트 대비 우수한 reference-based 성능.
- 추론 효율성 측면에서 기존 모델 대비 최소 50배의 latency 감소 및 약 2배 적은 메모리 사용.
- Reference-free 평가에서도 전용 MLLM과 유사한 성능 유지.
의의 및 한계
FlowTool은 자동 회귀적 추론 없이도 툴 파라미터를 생성할 수 있는 새로운 접근법을 제시하며, 이미지 리터치 작업의 효율성과 정확도를 동시에 향상시킨다. 특히, Diffusion Transformer와 흐름 매칭 기반 학습은 구조화된 연속 파라미터 공간을 생성하는 데 효과적임을 보여준다. 그러나 사용자 지시와 입력 이미지의 복잡성이 높아질 경우, 파라미터 생성의 정확도가 저하될 수 있다는 한계가 있다.
실용적 활용
FlowTool은 디지털 이미지 편집, 그래픽 디자인, 콘텐츠 제작 등에서 사용자 지시에 기반한 자동 툴 파라미터 조절을 지원할 수 있다. 특히, 빠른 추론 속도와 낮은 메모리 요구로 모바일 및 클라우드 기반 이미지 편집 애플리케이션에 적합하다.