General-Reasoner: Advancing LLM Reasoning Across All Domains

Xueguang Ma, Qian Liu, Dongfu Jiang, Ge Zhang, Zejun Ma, Wenhu Chen

arXiv:2505.14652 · 2026-08-15 공개 · arXiv · PDF

reinforcement-learning llm-reasoning mmlu gpqa mmlu-pro theoremqa zero-reinforcement-learning answer-verifier

Abstract

Reinforcement learning (RL) has recently demonstrated strong potential in enhancing the reasoning capabilities of large language models (LLMs). Particularly, the"Zero"reinforcement learning introduced by Deepseek-R1-Zero, enables direct RL training of base LLMs without relying on an intermediate supervised fine-tuning stage. Despite these advancements, current works for LLM reasoning mainly focus on mathematical and coding domains, largely due to data abundance and the ease of answer verification. This limits the applicability and generalization of such models to broader domains, where questions often have diverse answer representations, and data is more scarce. In this paper, we propose General-Reasoner, a novel training paradigm designed to enhance LLM reasoning capabilities across diverse domains. Our key contributions include: (1) constructing a large-scale, high-quality dataset of questions with verifiable answers curated by web crawling, covering a wide range of disciplines; and (2) developing a generative model-based answer verifier, which replaces traditional rule-based verification with the capability of chain-of-thought and context-awareness. We train a series of models and evaluate them on a wide range of datasets covering wide domains like physics, chemistry, finance, electronics etc. Our comprehensive evaluation across these 12 benchmarks (e.g. MMLU-Pro, GPQA, SuperGPQA, TheoremQA, BBEH and MATH AMC) demonstrates that General-Reasoner outperforms existing baseline methods, achieving robust and generalizable reasoning performance while maintaining superior effectiveness in mathematical reasoning tasks.

한국어 요약

한 줄 요약

General-Reasoner는 수학 외 다양한 분야에서 LLM의 추론 능력을 향상시키는 새로운 훈련 패러다임이다.

핵심 기여도

핵심 아이디어

기존 LLM 추론 훈련은 주로 수학 및 코딩 분야에 집중되어 있으며, 이는 데이터 풍부성과 답변 검증의 용이성 때문이었다. 그러나 이는 다른 분야로의 일반화를 제한한다. General-Reasoner는 두 가지 핵심 문제를 해결한다: (1) 다양한 분야의 검증 가능한 데이터 부족, (2) 규칙 기반 검증 방식의 제한성. 이를 위해 WebInstruct-verified라는 대규모 다분야 데이터셋을 구축하고, 생성 모델 기반의 General-Verifier를 도입하여 체인-오브-서스(Chain-of-Thought) 방식의 컨텍스트 인식 답변 검증을 가능하게 했다. 이는 Zero-RL 설정에서 훈련된 모델이 수학 외 분야에서도 뛰어난 추론 성능을 보이는 데 기여한다.

기술적 접근법

주요 결과

의의 및 한계

General-Reasoner는 LLM의 추론 능력을 수학 외 다양한 분야로 확장하는 데 기여하며, Zero-RL 설정에서의 훈련 효과를 입증했다. 생성 모델 기반 검증기는 규칙 기반 검증 방식의 제한성을 극복하고, 복잡한 답변 형식을 처리할 수 있다. 그러나 데이터셋은 웹 크롤링을 기반으로 하므로 일부 분야의 커버리지가 제한될 수 있으며, 생성 모델 기반 검증기의 오류 가능성도 존재한다. 또한, Zero-RL 설정은 훈련 안정성 측면에서 추가 연구가 필요할 수 있다.

실용적 활용

General-Reasoner는 과학, 금융, 사회과학 등 다양한 분야에서 복잡한 추론이 필요한 산업 및 연구 환경에 적용 가능하다. 특히, 데이터가 풍부하지 않은 분야에서의 모델 훈련 및 검증에 유용하며, 교육, 법률, 의료 분야의 전문가 시스템 개발에도 활용될 수 있다.