Energy-Guided Flow Matching

Haoyang Tong, Yu He, Fang Li, Lichen Ma, Jingling Fu, Dong Chen, Zhen Chen, Junshi Huang, Jie Cao

arXiv:2608.05811 · 2026-08-19 공개 · arXiv · PDF

flow-matching text-to-image image-generation high-resolution geneval imagenet dpg-bench fid-score

Abstract

Pixel-space generative models bypass lossy latent compression, yet necessitate joint learning of global structure and fine-grained details in a high-dimensional space. Standard flow matching interpolates noise toward a fixed clean-image endpoint, leaving the spectral evolution to be learned implicitly. In this paper, we introduce Energy-Guided Flow Matching(EG-FM) that explicitly models a coarse-to-fine generative trajectory by moving endpoint. Specifically, EG-FM replaces the fixed endpoint with a heat-kernel-filtered endpoint that evolves smoothly from low-frequency image to clean image. The fraction of high-frequency signal in moving endpoint is released by an image-specific energy-guided scheduling, leading to the re-targeting of velocity in flow matching. Our framework requires no adaptation of the backbone and training data, bringing negligible cost on the training and inference stages. In our experiment, EG-FM consistently achieves lower FID on the ImageNet class-conditional image generation task at 256 times 256 with fewer epochs, reaching an FID of 1.55 at 200 epochs and 1.45 at 600 epochs. We continue training the generation task on the setting of 512 times 512 resolution, yielding a FID of 1.58 after only 40 high-resolution adaptation epochs. Furthermore, we transfer EG-FM on text-to-image generation and achieve 0.85 on GenEval score and 83.9 on DPG-Bench. Code is available at https://github.com/ysng123/EG-FM.

한국어 요약

한 줄 요약

Energy-Guided Flow Matching(EG-FM)는 고주파 정보를 점진적으로 복원하는 생성 경로를 도입하여 FID 1.45를 달성한 픽셀 공간 생성 모델이다.

핵심 기여도

핵심 아이디어

기존 픽셀 공간 생성 모델은 고주파 세부 정보와 전역 구조를 동시에 학습해야 하므로 학습이 어려운 문제를 안고 있다. EG-FM은 이 문제를 해결하기 위해 생성 경로를 **coarse-to-fine** 방식으로 재설계한다. 구체적으로, 고정된 클린 이미지 엔드포인트 대신 **heat-kernel-filtered endpoint**를 사용하여, 이미지의 **low-frequency 구조부터 점진적으로 고주파 세부 정보를 복원**하도록 유도한다. 이는 **image-specific energy-guided scheduling**을 통해 각 이미지의 **spectral energy 분포에 따라 고주파 신호를 점진적으로 풀어내는 방식**이다. 이로 인해 생성 과정에서 **velocity vector가 재조정**되어, 모델이 먼저 전역 구조를 생성한 후 세부 텍스처를 생성할 수 있다. 이는 기존 flow matching에서 암묵적으로 학습해야 했던 주파수 진화를 명시적으로 모델링하는 핵심 아이디어이다.

기술적 접근법

주요 결과

의의 및 한계

EG-FM은 픽셀 공간 생성 모델에서 전역 구조와 세부 정보를 분리하여 학습하는 새로운 관점을 제시한다. 기존 flow matching이 암묵적으로 학습해야 했던 주파수 진화를 명시적으로 모델링함으로써, **더 빠른 수렴과 높은 생성 품질**을 달성할 수 있다. 또한, 기존 모델 구조나 학습 데이터를 수정하지 않아도 적용 가능하다는 점에서 실용적 가치가 높다. 그러나, **image-specific scheduling이 모든 이미지에 동일한 효과를 주는지**, **다양한 데이터셋에서의 일반화 능력**은 추가 연구가 필요하다. 또한, **high-resolution adaptation의 장기적 안정성**도 아직 명시되지 않았다.

실용적 활용

EG-FM은 고해상도 이미지 생성, 텍스트-이미지 생성, 이미지 편집 등 다양한 생성형 AI 분야에 적용 가능하다. 특히, 기존 픽셀 공간 모델을 최소한의 수정 없이 성능 향상이 필요한 산업 현장에서 유용하게 활용될 수 있다. 예를 들어, 디지털 콘텐츠 제작, 의료 이미지 생성, 자동차 디자인 등에서 빠른 수렴과 높은 품질의 생성이 요구되는 상황에 적합하다.