The Road Less Scheduled

Aaron Defazio, Xingyu Yang, Harsh Mehta, Konstantin Mishchenko, Ahmed Khaled, Ashok Cutkosky

arXiv:2405.15682 · 2026-07-27 공개 · arXiv · PDF

adamw optimization schedule-free learning-rate-scheduling iterate-averaging algoperf hyperparameter-free

Abstract

Existing learning rate schedules that do not require specification of the optimization stopping step T are greatly out-performed by learning rate schedules that depend on T. We propose an approach that avoids the need for this stopping time by eschewing the use of schedules entirely, while exhibiting state-of-the-art performance compared to schedules across a wide family of problems ranging from convex problems to large-scale deep learning problems. Our Schedule-Free approach introduces no additional hyper-parameters over standard optimizers with momentum. Our method is a direct consequence of a new theory we develop that unifies scheduling and iterate averaging. An open source implementation of our method is available at https://github.com/facebookresearch/schedule_free. Schedule-Free AdamW is the core algorithm behind our winning entry to the MLCommons 2024 AlgoPerf Algorithmic Efficiency Challenge Self-Tuning track.

한국어 요약

한 줄 요약

스케줄 없이 학습률을 조정하는 Schedule-Free 알고리즘이 기존 스케줄 기반 방법을 능가한다.

핵심 기여도

핵심 아이디어

기존 학습률 스케줄은 최적화 종료 시간 $ T $를 사전에 지정해야 하며, 이는 유연성과 실용성에 한계를 초래한다. 본 연구는 $ T $를 사전에 지정하지 않고도 Polyak-Ruppert (PR) 평균의 이론적 최적 수렴 속도를 유지하면서, 학습률 스케줄 기반 방법의 실용적 성능을 달성하는 새로운 접근법을 제시한다. 핵심 아이디어는 평균화(iterate averaging)와 학습률 스케줄을 이론적으로 통합하여, $ \beta $ 매개변수를 통해 PR 평균과 Primal 평균 사이의 보간(interpolation)을 수행하는 것이다. $ \beta = 0.9 $와 같은 표준 모멘텀 값이 실험적으로 효과적임을 보여주며, 이는 기존 EMA와 유사하지만, 나머지 기울기의 천천한 누적을 통해 안정성을 높인다.

기술적 접근법

주요 결과

의의 및 한계

실용적 활용