Training Deep Learning Models with Norm-Constrained LMOs

T. Pethick, Wanyun Xie, Kimon Antonakopoulos, Zhenyu Zhu, Antonio Silveti-Falls, V. Cevher

arXiv:2502.07529 · 2026-07-27 공개 · arXiv · PDF

optimization hyperparameter-transfer memory-efficient nanogpt stochastic-algorithms norm-constrained lmo scion

Abstract

In this work, we study optimization methods that leverage the linear minimization oracle (LMO) over a norm-ball. We propose a new stochastic family of algorithms that uses the LMO to adapt to the geometry of the problem and, perhaps surprisingly, show that they can be applied to unconstrained problems. The resulting update rule unifies several existing optimization methods under a single framework. Furthermore, we propose an explicit choice of norm for deep architectures, which, as a side benefit, leads to the transferability of hyperparameters across model sizes. Experimentally, we demonstrate significant speedups on nanoGPT training using our algorithm, Scion, without any reliance on Adam. The proposed method is memory-efficient, requiring only one set of model weights and one set of gradients, which can be stored in half-precision. The code is available at https://github.com/LIONS-EPFL/scion .

한국어 요약

한 줄 요약

Scion이라는 새로운 LMO 기반 최적화 알고리즘을 제안하여 nanoGPT 학습 속도를 향상시키고 메모리 효율성을 확보한다.

핵심 기여도

핵심 아이디어

기존 최적화 알고리즘은 학습 중 기하 구조를 동적으로 조정하지만, 이는 네트워크 구조에 대한 사전 정보를 무시한다. 본 연구는 LMO(Linear Minimization Oracle)를 사용하여 사전에 문제의 기하 구조에 맞게 최적화기를 설계하는 새로운 접근법을 제안한다. 특히, LMO는 노름-볼 내에서 선형 최소화를 수행하며, 이는 문제의 스케일에 관계없이 방향만 고려하는 특성을 가진다. 이 아이디어는 기존 Conditional Gradient 방법과 유사하지만, 벡터 ℓ∞-노름과 스펙트럼 노름과 같은 비표준 노름에서도 적용 가능하다는 점에서 차별화된다. 또한, Newton-Schultz 반복을 활용한 효율적인 LMO 계산을 제안하여, 실제 딥러닝 학습에 적용 가능성을 높인다.

기술적 접근법

주요 결과

의의 및 한계

실용적 활용

Scion 알고리즘은 Adam 없이도 빠른 학습 속도와 메모리 효율성을 제공하므로, 특히 메모리 제한이 있는 장치(예: 모바일, 임베디드)에서 딥러닝 모델 학습에 유용하게 활용될 수 있다. 또한, 하이퍼파라미터 전이 가능성은 다양한 모델 크기에서 동일한 최적화 전략을 적용할 수 있도록 지원한다.