The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning

Shivam Agarwal, Shivam Agarwal, Zimin Zhang, Lifan Yuan, Jiawei Han, Hao Peng

arXiv:2505.15134 · 2026-07-27 공개 · arXiv · PDF

reinforcement-learning llm-reasoning inference-time self-consistency logit-adjustment unlabeled-data entropy-minimization sci-code

Abstract

Entropy minimization (EM) trains the model to concentrate even more probability mass on its most confident outputs. We show that this simple objective alone, without any labeled data, can substantially improve large language models' (LLMs) performance on challenging math, physics, and coding tasks. We explore three approaches: (1) EM-FT minimizes token-level entropy similarly to instruction finetuning, but on unlabeled outputs drawn from the model; (2) EM-RL: reinforcement learning with negative entropy as the only reward to maximize; (3) EM-INF: inference-time logit adjustment to reduce entropy without any training data or parameter updates. On Qwen-7B, EM-RL, without any labeled data, achieves comparable or better performance than strong RL baselines such as GRPO and RLOO that are trained on 60K labeled examples. Furthermore, EM-INF enables Qwen-32B to match or exceed the performance of proprietary models like GPT-4o, Claude 3 Opus, and Gemini 1.5 Pro on the challenging SciCode benchmark, while being 3x more efficient than self-consistency and sequential refinement. Our findings reveal that many pretrained LLMs possess previously underappreciated reasoning capabilities that can be effectively elicited through entropy minimization alone, without any labeled data or even any parameter updates.

한국어 요약

한 줄 요약

엔트로피 최소화(EM)를 통해 라벨 없이도 LLM의 수학, 물리, 코딩 성능을 크게 향상시킬 수 있음을 보인 연구.

핵심 기여도

핵심 아이디어

엔트로피 최소화는 모델이 자신 있는 출력에 더 높은 확률을 집중시키도록 유도함으로써 추론 능력을 향상시키는 접근법이다. 기존 연구는 대부분 라벨 데이터나 파라미터 업데이트를 필요로 했지만, 본 연구는 **라벨 없이도** 모델의 내재된 추론 능력을 효과적으로 발휘할 수 있음을 보인다. 핵심 가정은 "모델이 자신 있을 때 더 정확하다"는 통찰이며, 이는 사전 학습된 LLM이 이미 충분한 능력을 갖추고 있다는 가정에 기반한다.

EM-FT는 미세조정과 유사한 방식으로, **토큰 수준 엔트로피**(Eqn. 2)를 최소화하는 손실 함수를 사용한다. EM-RL은 **음의 엔트로피**(Eqn. 1, 2)를 유일한 보상으로 사용하는 **REINFORCE 알고리즘**을 기반으로 한다. EM-INF는 추론 시 **로짓 조정**을 통해 **토큰 수준 엔트로피**(Eqn. 2)를 줄이는 방식으로, 파라미터 업데이트 없이도 추론 성능을 개선한다.

기술적 접근법

주요 결과

의의 및 한계

본 연구는 사전 학습된 LLM이 이미 강력한 추론 능력을 내재하고 있으며, 이를 **엔트로피 최소화만으로도 효과적으로 발휘**할 수 있음을 보인다. 이는 추론 성능 향상을 위한 라벨 데이터 의존성을 줄이고, 파라미터 업데이트 없이도 추론을 개선할 수 있는 실용적 방법을 제시한다.

그러나 EM은 모델의 **신뢰도가 정확성과 상관관계가 있을 때에만 효과적**이며, 인간 가치 정렬과 같은 작업에는 적합하지 않다. 또한, 사전 학습 모델이 해당 작업에 충분히 능력 있는 경우에만 EM이 효과적임을 보여주며, Qwen-2.5는 라벨 없이 EM만으로는 개선되지 않는 작업(개인적 가치 추론)에서는 실패함을 관찰했다.

실용적 활용

EM-FT, EM-RL, EM-INF는 추론 성능 향상을 위한 **라벨 데이터 없이도 사용 가능한 강력한 도구**로, 특히 수학, 물리, 코딩 등 **복잡한 추론 작업**에 적용 가능하다. Qwen-32B 기반 EM-INF는 **사전 학습된 모델의 추론 능력을 최대한 활용**하면서도 파라미터 업데이트 없이도 높은 효율성을 보임으로써, **대규모 모델의 추론 최적화**에 유용할 수 있다.