Flow Q-Learning

Seohong Park, Qiyang Li, Sergey Levine

arXiv:2502.02538 · 2026-07-27 공개 · arXiv · PDF

reinforcement-learning flow-matching policy-training offline-rl ogbench flow-q-learning action-distribution d4rl

Abstract

We present flow Q-learning (FQL), a simple and performant offline reinforcement learning (RL) method that leverages an expressive flow-matching policy to model arbitrarily complex action distributions in data. Training a flow policy with RL is a tricky problem, due to the iterative nature of the action generation process. We address this challenge by training an expressive one-step policy with RL, rather than directly guiding an iterative flow policy to maximize values. This way, we can completely avoid unstable recursive backpropagation, eliminate costly iterative action generation at test time, yet still mostly maintain expressivity. We experimentally show that FQL leads to strong performance across 73 challenging state- and pixel-based OGBench and D4RL tasks in offline RL and offline-to-online RL. Project page: https://seohong.me/projects/fql/

한국어 요약

한 줄 요약

Flow Q-Learning(FQL)은 복잡한 행동 분포를 모델링하는 flow policy를 활용한 오프라인 강화학습 방법으로, 73개의 OGBench 및 D4RL 태스크에서 우수한 성능을 보인다.

핵심 기여도

핵심 아이디어

기존 flow 또는 diffusion policy를 사용한 오프라인 RL에서는 반복적 생성 과정으로 인해 backpropagation이 불안정하고, 테스트 시 비용이 높은 문제가 있었다. FQL은 이 문제를 해결하기 위해 flow policy를 행동 복제(behavior cloning, BC)로만 학습하고, 별도의 one-step policy를 Q-learning으로 학습한다. 이 one-step policy는 flow policy로부터 지식 증류(distillation)를 통해 학습되며, 이는 value 최대화를 flow policy로부터 분리하여 불안정한 반복 과정을 회피하고, 테스트 시 반복적 flow step을 제거한다. 이 접근법은 flow model의 표현력은 유지하면서도 학습 및 추론 효율성을 동시에 달성한다.

기술적 접근법

주요 결과

의의 및 한계

FQL은 복잡한 행동 분포를 모델링하면서도, 반복적 생성 과정을 제거함으로써 테스트 시 비용을 절감하고, 불안정한 backpropagation을 회피하는 점에서 기존 flow 및 diffusion policy 기반 방법보다 실용적이다. 또한, one-step policy를 통해 Q-learning과의 호환성을 높여, 기존 actor-critic 프레임워크에 쉽게 통합할 수 있다. 그러나 flow policy가 데이터 분포를 완벽히 모델링하지 못하는 경우, distillation 과정에서 정보 손실이 발생할 수 있다. 또한, BC coefficient $ \alpha $는 환경에 따라 민감하게 조정되어야 하며, 이는 추가적인 하이퍼파라미터 튜닝을 필요로 한다.

실용적 활용

FQL은 로봇 제어, 자율 주행, 게임 AI 등 복잡한 행동 분포를 요구하는 오프라인 강화학습 문제에 적용 가능하다. 특히, 테스트 시 반복적 생성 과정이 필요 없는 점에서 실시간 성능이 중요한 산업 현장에서 유리하며, online fine-tuning을 통해 환경 변화에 빠르게 적응할 수 있다.