Aya Model: An Instruction Finetuned Open-Access Multilingual Language Model

A. Ustun, Viraat Aryabumi, Zheng-Xin Yong, Wei-Yin Ko, Daniel D'souza, Gbemileke Onilude, Neel Bhandari, Shivalika Singh, Hui-Lee Ooi, Amr Kayid, Freddie Vargus, Phil Blunsom, Shayne Longpre, Niklas Muennighoff, Marzieh Fadaee, Julia Kreutzer, Sara Hooker

arXiv:2402.07827 · 2026-07-27 공개 · arXiv · PDF

benchmarking model-evaluation llm-safety low-resource-languages instruction-finetuning data-pruning multilingual-model aya-model

Abstract

Recent breakthroughs in large language models (LLMs) have centered around a handful of data-rich languages. What does it take to broaden access to breakthroughs beyond first-class citizen languages? Our work introduces Aya, a massively multilingual generative language model that follows instructions in 101 languages of which over 50% are considered as lower-resourced. Aya outperforms mT0 and BLOOMZ on the majority of tasks while covering double the number of languages. We introduce extensive new evaluation suites that broaden the state-of-art for multilingual eval across 99 languages -- including discriminative and generative tasks, human evaluation, and simulated win rates that cover both held-out tasks and in-distribution performance. Furthermore, we conduct detailed investigations on the optimal finetuning mixture composition, data pruning, as well as the toxicity, bias, and safety of our models. We open-source our instruction datasets and our model at https://hf.co/CohereForAI/aya-101

한국어 요약

한 줄 요약

Aya는 101개 언어(51개는 저자원 언어)에서 지시사항을 따르는 오픈소스 멀티링구 모델로, mT0와 BLOOMZ보다 대부분의 작업에서 성능이 우수하다.

핵심 기여도

핵심 아이디어

Aya는 기존 멀티링구 모델이 대부분 고자원 언어에 집중된 반면, 저자원 언어를 포함한 다국어 지시사항 수행 능력을 확장하는 데 초점을 맞췄다. 이는 기존 IFT 데이터셋이 영어에 치우쳐 있다는 문제를 해결하기 위한 접근이다. Aya는 xP3 데이터셋을 기반으로 101개 언어로 확장한 203M 데이터를 사용했으며, 데이터 가중치 조정과 트리밍을 통해 언어 분포를 균형 있게 조정했다. 특히, 인간 주석을 기반으로 영어 인스턴스를 19.66%, 다국어 인스턴스를 18.25% 제거함으로써 언어 편향을 줄였다. 이는 기존 모델(mT0, BLOOMZ)이 다루지 못한 언어 범위와 평가 범위를 확장하는 데 기여한다.

기술적 접근법

주요 결과

의의 및 한계

Aya는 기존 멀티링구 모델이 다루지 못한 저자원 언어를 포함한 다국어 지시 수행 능력을 확장한 점에서 학술적·실용적 의의가 있다. 특히, 101개 언어를 지원하며, 이 중 51개는 저자원 언어로, 언어 불평등 문제를 완화할 수 있다. 또한, 99개 언어에 걸쳐 평가 범위를 확장함으로써 기존 연구의 한계를 극복했다. 그러나, 7,000개 언어 중 93%는 여전히 모델 학습에 포함되지 않아 언어 다각성 확보는 여전히 어려운 과제이다. 또한, 13B 파라미터 규모는 소비자용 하드웨어에서 실행이 어려워 압축 기술(퀀티제이션, 트리밍)이 필요하다.

실용적 활용

Aya는 저자원 언어 사용자에게 더 나은 품질의 언어 모델을 제공할 수 있으며, 글로벌 기업이나 교육 기관에서 다국어 지원 서비스를 개선하는 데 활용될 수 있다. 또한, 언어 평가 및 안전성 검증이 필요한 연구 분야에서 평가 기준 확장에 기여할 수 있다.