RedPajama: an Open Dataset for Training Large Language Models

Maurice Weber, Dan Fu, Quentin Anthony, Yonatan Oren, Shane Adams, Anton Alexandrov, Xiaozhong Lyu, Huu Nguyen, Xiaozhe Yao, Virginia Adams, Ben Athiwaratkun, Rahul Chalamala, Kezhen Chen, Max Ryabinin, Tri Dao, Percy Liang, Christopher R'e, Irina Rish, Ce Zhang

arXiv:2411.12372 · 2026-07-27 공개 · arXiv · PDF

large-language-models open-source dataset-curation decoder-only metadata redpajama llama-reproduction web-text

Abstract

Large language models are increasingly becoming a cornerstone technology in artificial intelligence, the sciences, and society as a whole, yet the optimal strategies for dataset composition and filtering remain largely elusive. Many of the top-performing models lack transparency in their dataset curation and model development processes, posing an obstacle to the development of fully open language models. In this paper, we identify three core data-related challenges that must be addressed to advance open-source language models. These include (1) transparency in model development, including the data curation process, (2) access to large quantities of high-quality data, and (3) availability of artifacts and metadata for dataset curation and analysis. To address these challenges, we release RedPajama-V1, an open reproduction of the LLaMA training dataset. In addition, we release RedPajama-V2, a massive web-only dataset consisting of raw, unfiltered text data together with quality signals and metadata. Together, the RedPajama datasets comprise over 100 trillion tokens spanning multiple domains and with their quality signals facilitate the filtering of data, aiming to inspire the development of numerous new datasets. To date, these datasets have already been used in the training of strong language models used in production, such as Snowflake Arctic, Salesforce's XGen and AI2's OLMo. To provide insight into the quality of RedPajama, we present a series of analyses and ablation studies with decoder-only language models with up to 1.6B parameters. Our findings demonstrate how quality signals for web data can be effectively leveraged to curate high-quality subsets of the dataset, underscoring the potential of RedPajama to advance the development of transparent and high-performing language models at scale.

한국어 요약

한 줄 요약

RedPajama는 100조 토큰 이상의 오픈 소스 대형 언어 모델 학습 데이터셋으로, 투명성과 품질 신호를 통해 모델 성능을 개선한다.

핵심 기여도

핵심 아이디어

RedPajama는 대형 언어 모델의 훈련 데이터셋 개발에서 투명성, 규모, 다용도성을 핵심 원칙으로 삼았다. 기존 모델들이 데이터셋 구성과 필터링 전략을 공개하지 않는 문제를 해결하기 위해, RedPajama는 데이터셋 생성 과정을 완전히 공개하고, 품질 신호를 함께 제공함으로써 사용자가 데이터셋을 선택적 필터링할 수 있도록 지원한다. 특히, RedPajama-V2는 웹 기반 원시 데이터에 46개의 품질 신호를 부여하여, 다양한 필터링 전략을 실험할 수 있는 기반을 제공한다. 이는 기존의 C4, Gopher 규칙과 비교해 더 유연한 데이터셋 구성이 가능하다는 점에서 혁신적이다.

기술적 접근법

주요 결과

의의 및 한계

RedPajama는 대형 언어 모델 개발의 투명성과 재현성을 높이는 데 기여하며, 다양한 필터링 전략을 실험할 수 있는 기반 데이터셋으로 활용 가능하다. 특히, 품질 신호를 기반으로 한 데이터셋 필터링이 성능에 직접적인 영향을 미친다는 점에서 학술적 가치가 크다. 그러나 본 연구는 1.6B 파라미터 이하의 상대적으로 작은 모델만 사용했으며, 대규모 모델에서의 일반화 가능성은 추가 연구가 필요하다. 또한, 개인 정보 유출 가능성이나 데이터셋 오염 분석은 수행되지 않았다는 한계점이 있다.

실용적 활용

RedPajama는 Snowflake Arctic, Salesforce XGen, AI2 OLMo 등 실제 산업용 언어 모델의 훈련에 이미 활용되고 있으며, 향후 데이터셋 필터링 및 조합 전략 연구에 폭넓게 사용될 수 있다. 특히, 품질 신호를 기반으로 한 데이터셋 필터링은 클라우드 기반 대규모 모델 훈련 시 데이터 품질 관리를 위한 실용적 도구로 활용 가능하다.