Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research

Luca Soldaini, Rodney Kinney, Akshita Bhagia, Dustin Schwenk, David Atkinson, Russell Authur, Ben Bogin, Khyathi Raghavi Chandu, Jennifer Dumas, Yanai Elazar, Valentin Hofmann, A. Jha, Sachin Kumar, L. Lucy, Xinxi Lyu, Nathan Lambert, Ian Magnusson, Jacob Daniel Morrison, Niklas Muennighoff, Aakanksha Naik, Crystal Nam, Matthew E. Peters, Abhilasha Ravichander, Kyle Richardson, Shannon Zejiang Shen, Emma Strubell, Nishant Subramani, Oyvind Tafjord, Pete Walsh, Luke S. Zettlemoyer, Noah A. Smith, Hanna Hajishirzi, Iz Beltagy, Dirk Groeneveld, Jesse Dodge, Kyle Lo

arXiv:2402.00159 · 2026-07-27 공개 · arXiv · PDF

language-models open-source tokenization pretraining-data large-scale-data model-pretraining web-content corpus-curation

Abstract

Information about pretraining corpora used to train the current best-performing language models is seldom discussed: commercial models rarely detail their data, and even open models are often released without accompanying training data or recipes to reproduce them. As a result, it is challenging to conduct and advance scientific research on language modeling, such as understanding how training data impacts model capabilities and limitations. To facilitate scientific research on language model pretraining, we curate and release Dolma, a three-trillion-token English corpus, built from a diverse mixture of web content, scientific papers, code, public-domain books, social media, and encyclopedic materials. We extensively document Dolma, including its design principles, details about its construction, and a summary of its contents. We present analyses and experimental results on intermediate states of Dolma to share what we have learned about important data curation practices. Finally, we open-source our data curation toolkit to enable reproduction of our work as well as support further research in large-scale data curation.

한국어 요약

한 줄 요약

Dolma는 3조 토큰 규모의 영어 사전 학습 데이터셋으로, OLMo 모델 학습에 사용된 개방형 언어 모델 연구를 위한 기반 자료이다.

핵심 기여도

핵심 아이디어

현재 최고 성능을 내는 언어 모델의 사전 학습 데이터셋 정보는 대부분 비공개이거나 불완전하여, 연구 재현 및 데이터 영향 분석이 어렵다. 이에 따라 Dolma는 데이터 투명성과 재현성을 목표로, 공개 가능한 소스에서 3조 토큰을 수집하고, 다양한 큐레이션 단계를 적용하여 고품질 데이터셋을 구축했다. 특히, 품질 필터(Quality Filter), 언어 필터(Language Filter), 중복 제거(Deduplication), 벤치마크 오염 제거(Benchmark Decontamination) 등의 단계를 통해 데이터셋의 신뢰성을 높였다. 또한, 데이터셋의 구성이 모델 성능에 미치는 영향을 12개의 QA, 상식, 추론 태스크에서 실증적으로 분석했다.

기술적 접근법

주요 결과

의의 및 한계

Dolma는 언어 모델 사전 학습 연구에서 데이터 투명성과 재현성을 촉진하는 중요한 기초 자료로, 다양한 큐레이션 전략의 효과를 실증적으로 평가할 수 있는 기회를 제공한다. 특히, Dolma Toolkit은 데이터 큐레이션 파이프라인의 표준화와 개선에 기여할 수 있다. 그러나 Dolma는 영어 중심이며, 비영어 언어 모델 연구에는 한계가 있다. 또한, 데이터 소스의 공개성과 접근성에 따라 일부 데이터는 제외되거나 제한될 수 있다.

실용적 활용

Dolma는 언어 모델 개발자, 연구자, 데이터 큐레이션 엔지니어에게 데이터셋 개선, 모델 성능 분석, 큐레이션 전략 연구에 활용 가능하다. 특히, OLMo와 같은 개방형 언어 모델의 학습 및 평가에 적합하며, 데이터 품질과 구성이 모델 성능에 미치는 영향을 연구하는 데 유용하다.