OpenMathInstruct-2: Accelerating AI for Math with Massive Open-Source Instruction Data

Shubham Toshniwal, Wei Du, Ivan Moshkov, B. Kisačanin, Alexan Ayrapetyan, Igor Gitman

arXiv:2410.01560 · 2026-07-27 공개 · arXiv · PDF

model-evaluation open-source math-reasoning llm-finetuning data-synthesis dataset-generation instruction-data llama3-1

Abstract

Mathematical reasoning continues to be a critical challenge in large language model (LLM) development with significant interest. However, most of the cutting-edge progress in mathematical reasoning with LLMs has become \emph{closed-source} due to lack of access to training data. This lack of data access limits researchers from understanding the impact of different choices for synthesizing and utilizing the data. With the goal of creating a high-quality finetuning (SFT) dataset for math reasoning, we conduct careful ablation experiments on data synthesis using the recently released \texttt{Llama3.1} family of models. Our experiments show that: (a) solution format matters, with excessively verbose solutions proving detrimental to SFT performance, (b) data generated by a strong teacher outperforms equally-sized data generated by a weak student model, (c) SFT is robust to low-quality solutions, allowing for imprecise data filtering, and (d) question diversity is crucial for achieving data scaling gains. Based on these insights, we create the OpenMathInstruct-2 dataset, which consists of 14M question-solution pairs ($\approx$ 600K unique questions), making it nearly eight times larger than the previous largest open-source math reasoning dataset. Finetuning the \texttt{Llama-3.1-8B-Base} using OpenMathInstruct-2 outperforms \texttt{Llama3.1-8B-Instruct} on MATH by an absolute 15.9\% (51.9\% $\rightarrow$ 67.8\%). Finally, to accelerate the open-source efforts, we release the code, the finetuned models, and the OpenMathInstruct-2 dataset under a commercially permissive license.

한국어 요약

한 줄 요약

OpenMathInstruct-2는 Llama3.1을 활용해 생성된 14M개의 수학 문제-해답 쌍을 포함한 오픈소스 학습 데이터셋으로, Llama3.1-8B-Base 모델을 fine-tuning한 결과 MATH 데이터셋에서 15.9% 성능 향상을 달성했다.

핵심 기여도

핵심 아이디어

수학 추론 능력을 향상시키기 위한 학습 데이터셋은 대부분 비공개이거나 제한된 라이선스로 인해 연구자들이 데이터 구성과 합성 방식의 영향을 이해하기 어려웠다. 본 연구는 Llama3.1 모델을 활용해 **고품질 수학 학습 데이터를 생성**하고, 이를 기반으로 **instruction tuning 데이터셋인 OpenMathInstruct-2를 개발**했다. 핵심 통찰은 다음과 같다:

기술적 접근법

  • **데이터 생성**: Llama3.1-405B-Instruct 모델을 teacher로 사용하여 MATH와 GSM8K 문제에 대한 해답을 합성하고, 새로운 문제-해답 쌍을 생성.
  • **해답 형식**: 기존 Llama CoT 대비 40% 짧은 해답 길이를 유지하면서도 3.9% 더 높은 성능을 보이는 **개선된 CoT 형식** 사용.
  • **SFT 학습**: Llama3.1-8B-Base 모델을 4 에포크, 배치 크기 256, 학습률 5e-6, weight decay 1e-2로 학습.
  • **데이터 정제**: lm-sys 파이프라인과 수작업 검토를 통해 **테스트셋 오염 방지**.
  • **성능 평가**: GPT-4o를 사용한 자동 평가와 256개 샘플에 대한 다수결 투표를 통해 정확도 측정.

주요 결과

  • **MATH 데이터셋에서 OpenMath2-Llama3.1-8B 모델**이 Llama3.1-8B-Instruct 대비 **15.9% 절대 개선** (51.9% → 67.8%).
  • **OpenMath2-Llama3.1-70B 모델**은 MATH에서 **71.9% 정확도** 달성, Llama3.1-70B-Instruct 대비 **3.9% 개선**.
  • **질문 다양성 증가**(1K → 6.5K)로 MATH 검증 집합에서 **10.5% 절대 개선**.
  • **20%의 저품질 데이터 포함 시에도 SFT 성능 유지**.

의의 및 한계

OpenMathInstruct-2는 수학 추론 분야에서 **가장 큰 규모의 오픈소스 학습 데이터셋**으로, 연구자들이 데이터 구성과 합성 방식의 영향을 자유롭게 분석할 수 있도록 기여한다. 또한, 학습 데이터 생성 과정에서 **강력한 teacher 모델의 활용**과 **해답 형식 최적화**를 통해 **고품질 학습 데이터를 생성**할 수 있음을 보여준다. 그러나 본 연구는 **GSM8K, MATH 등 기존 문제 기반의 합성 데이터**를 사용하기 때문에, **완전히 새로운 문제 유형의 생성 가능성**은 명시되지 않았다. 또한, **비공개 모델**(예: GPT-4o)을 사용한 평가가 포함되어 있어, **오픈소스 모델 간 비교의 공정성**에 대한 논란이 있을 수 있다.

실용적 활용

OpenMathInstruct-2는 수학 교육, AI 모델 학습, 학술 연구 등에서 **고품질 수학 학습 데이터가 필요한 모든 분야**에 적용 가능하다. 특히, **Llama3.1과 같은 오픈소스 모델의 학습 데이터 확장** 및 **instruction tuning 최적화**에 활용할 수 있으며, **비상업적 및 상업적 용도 모두 가능**하다는 라이선스 특성 덕분에 **산업 및 연구 기관에서 폭넓게 사용**될 수 있다.