Aya Dataset: An Open-Access Collection for Multilingual Instruction Tuning

Shivalika Singh, Freddie Vargus, Daniel Dsouza, Börje F. Karlsson, Abinaya Mahendiran, Wei-Yin Ko, Herumb Shandilya, J. Patel, Deividas Mataciunas, Laura O'Mahony, Mike Zhang, Ramith Hettiarachchi, Joseph Wilson, Marina Machado, L. S. Moura, Dominik Krzemi'nski, Hakimeh Fadaei, Irem Ergun, Ifeoma Okoh, Aisha Alaagib, Oshan Mudannayake, Zaid Alyafeai, Minh Chien Vu, Sebastian Ruder, Surya Guthikonda, Emad A. Alghamdi, Sebastian Gehrmann, Niklas Muennighoff, M. Bartolo, Julia Kreutzer, A. Ustun, Marzieh Fadaee, Sara Hooker

arXiv:2402.06619 · 2026-07-27 공개 · arXiv · PDF

llm instruction-tuning dataset multilingual natural-language-processing language-resources evaluation-suite annotation-platform

Abstract

Datasets are foundational to many breakthroughs in modern artificial intelligence. Many recent achievements in the space of natural language processing (NLP) can be attributed to the finetuning of pre-trained models on a diverse set of tasks that enables a large language model (LLM) to respond to instructions. Instruction fine-tuning (IFT) requires specifically constructed and annotated datasets. However, existing datasets are almost all in the English language. In this work, our primary goal is to bridge the language gap by building a human-curated instruction-following dataset spanning 65 languages. We worked with fluent speakers of languages from around the world to collect natural instances of instructions and completions. Furthermore, we create the most extensive multilingual collection to date, comprising 513 million instances through templating and translating existing datasets across 114 languages. In total, we contribute four key resources: we develop and open-source the Aya Annotation Platform, the Aya Dataset, the Aya Collection, and the Aya Evaluation Suite. The Aya initiative also serves as a valuable case study in participatory research, involving collaborators from 119 countries. We see this as a valuable framework for future research collaborations that aim to bridge gaps in resources.

한국어 요약

한 줄 요약

Aya Dataset은 65개 언어의 20만 개 이상의 인간 주도 인스트럭션-완료 쌍을 포함한 최대 규모의 오픈소스 다국어 학습 데이터셋이다.

핵심 기여도

핵심 아이디어

기존의 인스트럭션 퍼파인튜닝(IFT) 데이터셋은 대부분 영어에 집중되어 있으며, 이는 언어 간 불균형을 심화시키는 주요 원인이다. Aya Dataset은 119개 국가에서 2,997명의 모국어 사용자와 협력하여 65개 언어의 자연스러운 인스트럭션-완료 쌍을 수집함으로써 이 문제를 해결하려는 시도이다. 특히, Aya는 단순히 기계 번역이나 자동화된 데이터셋 생성 방식이 아닌, 인간 주도의 데이터 수집 프로세스를 중점적으로 설계하여 번역 편향과 문화적 맥락의 누락을 최소화했다. Aya Annotation Platform은 사용자 친화적인 UI를 통해 글로벌 참여자들이 데이터를 쉽게 기여할 수 있도록 지원하며, 이는 오픈소스 및 참여형 과학 프로젝트의 최선 사례를 반영한 설계이다.

기술적 접근법

주요 결과

의의 및 한계

Aya Dataset은 다국어 NLP 연구에서의 언어 불균형 문제를 해결하기 위한 중요한 자산으로, 특히 저자원 언어 사용자들이 기술 접근성을 높이는 데 기여할 수 있다. 또한, 참여형 연구와 오픈소스 과학의 모범 사례로, 사회언어학자, 인류학자 등 다양한 학문 분야와의 협업 가능성을 열어준다. 그러나 Aya는 여전히 자원이 제한된 언어에 대한 데이터 수집이 부족하며, 자동화된 데이터셋 생성 방식에 비해 확장성이 낮다는 한계가 있다.

실용적 활용

Aya Dataset은 다국어 대형 언어 모델의 학습 및 평가에 활용될 수 있으며, 특히 저자원 언어 사용자들이 AI 기술의 혜택을 받을 수 있도록 지원한다. 또한, 글로벌 협력 기반의 데이터셋 개발 프로세스는 언어 기술 개발과 사회적 책임을 결합한 연구 프로젝트에 모범 사례를 제공한다.