FlowTool: Controlling Tool Parameter in Image Retouching via Flow Matching

Thanh-Long V. Le, Steven Walton, Seunghyun Yoon, Branislav Kveton, Trung Bui, Eunho Yang, Viet Lai

arXiv:2609.35673 · 2026-09-29 공개 · arXiv · PDF

vision-language flow-matching diffusion-transformer rectified-flow conditional-generation tool-based-editing image-retouching mmart-bench

Abstract

Tool-based image editing (image retouching) is commonly formulated with autoregressive multimodal large language models (MLLMs) that sequentially generate reasoning, tool selections, and parameter values. In this work, we present a novel approach to tool-based image editing by framing the task as a flow matching problem. We introduce FlowTool, a framework that directly models the distribution of high-quality tool parameters conditioned on the input image and user instruction using conditional rectified flow. FlowTool combines a vision-language model backbone for multimodal understanding with a Diffusion Transformer parameter generator that transforms Gaussian noise into an editing plan. We train FlowTool with a two-stage supervised flow-matching curriculum, followed by reward-based post-training. Across MMArt-Bench, FlowTool-Eval, ArtEdit-Bench, and MIT-Adobe5K, FlowTool achieves significantly stronger reference-based performance than specialized MLLM editing agents and proprietary MLLMs, while remaining competitive with proprietary models under reference-free evaluation. Moreover, FlowTool significantly improves inference efficiency, reducing latency by at least 50times while requiring nearly 2times less memory than the compared baselines. These results demonstrate that tool-based image editing can be effectively modeled as conditional generation over structured continuous editing parameters, without autoregressive reasoning.

한국어 요약

한 줄 요약

FlowTool은 이미지 리터치 작업에서 툴 파라미터를 조절하기 위해 흐름 매칭을 활용한 새로운 프레임워크이다.

핵심 기여도

핵심 아이디어

기존 이미지 리터치는 자동 회귀적 MLLM을 사용해 툴 선택과 파라미터 생성을 순차적으로 처리했으나, 이는 추론 단계가 복잡하고 비효율적이다. FlowTool은 이러한 문제를 해결하기 위해 툴 파라미터 생성을 흐름 매칭 문제로 재구성한다. 이 접근법은 조건부 정규화 흐름을 통해 이미지와 사용자 지시를 기반으로 연속적인 편집 파라미터를 생성함으로써, 자동 회귀적 추론 없이도 효과적인 편집 계획을 도출할 수 있다. Diffusion Transformer는 노이즈를 편집 계획으로 변환하는 역할을 하며, 이는 기존 방식보다 구조화된 파라미터 공간을 직접 생성할 수 있다는 장점을 가진다.

기술적 접근법

주요 결과

의의 및 한계

FlowTool은 자동 회귀적 추론 없이도 툴 파라미터를 생성할 수 있는 새로운 접근법을 제시하며, 이미지 리터치 작업의 효율성과 정확도를 동시에 향상시킨다. 특히, Diffusion Transformer와 흐름 매칭 기반 학습은 구조화된 연속 파라미터 공간을 생성하는 데 효과적임을 보여준다. 그러나 사용자 지시와 입력 이미지의 복잡성이 높아질 경우, 파라미터 생성의 정확도가 저하될 수 있다는 한계가 있다.

실용적 활용

FlowTool은 디지털 이미지 편집, 그래픽 디자인, 콘텐츠 제작 등에서 사용자 지시에 기반한 자동 툴 파라미터 조절을 지원할 수 있다. 특히, 빠른 추론 속도와 낮은 메모리 요구로 모바일 및 클라우드 기반 이미지 편집 애플리케이션에 적합하다.