reinforcement-learning llm-reasoning mmlu gpqa mmlu-pro theoremqa zero-reinforcement-learning answer-verifier
Abstract
Reinforcement learning (RL) has recently demonstrated strong potential in enhancing the reasoning capabilities of large language models (LLMs). Particularly, the"Zero"reinforcement learning introduced by Deepseek-R1-Zero, enables direct RL training of base LLMs without relying on an intermediate supervised fine-tuning stage. Despite these advancements, current works for LLM reasoning mainly focus on mathematical and coding domains, largely due to data abundance and the ease of answer verification. This limits the applicability and generalization of such models to broader domains, where questions often have diverse answer representations, and data is more scarce. In this paper, we propose General-Reasoner, a novel training paradigm designed to enhance LLM reasoning capabilities across diverse domains. Our key contributions include: (1) constructing a large-scale, high-quality dataset of questions with verifiable answers curated by web crawling, covering a wide range of disciplines; and (2) developing a generative model-based answer verifier, which replaces traditional rule-based verification with the capability of chain-of-thought and context-awareness. We train a series of models and evaluate them on a wide range of datasets covering wide domains like physics, chemistry, finance, electronics etc. Our comprehensive evaluation across these 12 benchmarks (e.g. MMLU-Pro, GPQA, SuperGPQA, TheoremQA, BBEH and MATH AMC) demonstrates that General-Reasoner outperforms existing baseline methods, achieving robust and generalizable reasoning performance while maintaining superior effectiveness in mathematical reasoning tasks.
한국어 요약
한 줄 요약
General-Reasoner는 수학 외 다양한 분야에서 LLM의 추론 능력을 향상시키는 새로운 훈련 패러다임이다.
핵심 기여도
- WebInstruct-verified라는 230K개의 고질량, 다분야 추론 질문 데이터셋을 구축.
- General-Verifier라는 생성 모델 기반의 답변 검증기(1.5B 파라미터)를 개발, 기존 규칙 기반 검증 방식 대체.
- Zero-RL 설정에서 훈련한 General-Reasoner 모델이 MMLU-Pro, SuperGPQA 등 12개 벤치마크에서 기존 베이스라인 대비 약 10% 성능 향상.
- 수학 분야에서도 SimpleRL과 유사하거나 더 우수한 성능 달성.
핵심 아이디어
기존 LLM 추론 훈련은 주로 수학 및 코딩 분야에 집중되어 있으며, 이는 데이터 풍부성과 답변 검증의 용이성 때문이었다. 그러나 이는 다른 분야로의 일반화를 제한한다. General-Reasoner는 두 가지 핵심 문제를 해결한다: (1) 다양한 분야의 검증 가능한 데이터 부족, (2) 규칙 기반 검증 방식의 제한성. 이를 위해 WebInstruct-verified라는 대규모 다분야 데이터셋을 구축하고, 생성 모델 기반의 General-Verifier를 도입하여 체인-오브-서스(Chain-of-Thought) 방식의 컨텍스트 인식 답변 검증을 가능하게 했다. 이는 Zero-RL 설정에서 훈련된 모델이 수학 외 분야에서도 뛰어난 추론 성능을 보이는 데 기여한다.
기술적 접근법
- **WebInstruct-verified 데이터셋**: 웹 크롤링을 통해 230K개의 다분야 추론 질문을 수집 및 필터링.
- **General-Verifier**: 1.5B 파라미터의 생성 모델 기반 검증기. 체인-오브-서스 방식으로 컨텍스트 인식 답변 검증.
- **Zero-RL 훈련**: 기존의 중간 단계인 감독 훈련 없이, GRPO 알고리즘을 사용해 기반 LLM을 직접 훈련.
- **모델**: General-Reasoner-Qw3-14B가 GPT-4o 수준의 성능을 달성.
- **평가 프레임워크**: simple-evals 2를 사용해 GPT4o와 답변 동등성 검증.
주요 결과
- MMLU-Pro에서 10% (베이스라인 대비 +10%) 성능 향상.
- SuperGPQA에서도 유사한 수준의 개선.
- 수학 벤치마크(MATH-500, GSM8K 등)에서도 SimpleRL과 유사하거나 더 우수한 성능.
- General-Reasoner-Qw3-14B가 GPT-4o 수준의 성능을 달성.
의의 및 한계
General-Reasoner는 LLM의 추론 능력을 수학 외 다양한 분야로 확장하는 데 기여하며, Zero-RL 설정에서의 훈련 효과를 입증했다. 생성 모델 기반 검증기는 규칙 기반 검증 방식의 제한성을 극복하고, 복잡한 답변 형식을 처리할 수 있다. 그러나 데이터셋은 웹 크롤링을 기반으로 하므로 일부 분야의 커버리지가 제한될 수 있으며, 생성 모델 기반 검증기의 오류 가능성도 존재한다. 또한, Zero-RL 설정은 훈련 안정성 측면에서 추가 연구가 필요할 수 있다.
실용적 활용
General-Reasoner는 과학, 금융, 사회과학 등 다양한 분야에서 복잡한 추론이 필요한 산업 및 연구 환경에 적용 가능하다. 특히, 데이터가 풍부하지 않은 분야에서의 모델 훈련 및 검증에 유용하며, 교육, 법률, 의료 분야의 전문가 시스템 개발에도 활용될 수 있다.