How do Large Language Models Handle Multilingualism?

Yiran Zhao, Wenxuan Zhang, Guizhen Chen, Kenji Kawaguchi, Lidong Bing

arXiv:2402.18815 · 2026-07-27 공개 · arXiv · PDF

large-language-models fine-tuning self-attention low-resource-languages neuron-detection language-specific-neurons high-resource-languages multilingualism

Abstract

Large language models (LLMs) have demonstrated impressive capabilities across diverse languages. This study explores how LLMs handle multilingualism. Based on observed language ratio shifts among layers and the relationships between network structures and certain capabilities, we hypothesize the LLM's multilingual workflow ($\texttt{MWork}$): LLMs initially understand the query, converting multilingual inputs into English for task-solving. In the intermediate layers, they employ English for thinking and incorporate multilingual knowledge with self-attention and feed-forward structures, respectively. In the final layers, LLMs generate responses aligned with the original language of the query. To verify $\texttt{MWork}$, we introduce Parallel Language-specific Neuron Detection ($\texttt{PLND}$) to identify activated neurons for inputs in different languages without any labeled data. Using $\texttt{PLND}$, we validate $\texttt{MWork}$ through extensive experiments involving the deactivation of language-specific neurons across various layers and structures. Moreover, $\texttt{MWork}$ allows fine-tuning of language-specific neurons with a small dataset, enhancing multilingual abilities in a specific language without compromising others. This approach results in an average improvement of $3.6\%$ for high-resource languages and $2.3\%$ for low-resource languages across all tasks with just $400$ documents.

한국어 요약

한 줄 요약

LLMs는 다국어 처리를 영어로 변환-이해-생성하는 3단계 워크플로우(MWork)를 통해 수행하며, PLND를 통해 검증하고 언어별 정밀 조정이 가능하다.

핵심 기여도

핵심 아이디어

기존 연구는 영어 중심으로 LLM의 다국어 처리 메커니즘을 해석했으나, 본 연구는 LLM이 영어를 중간 언어로 사용하는 3단계 워크플로우(MWork)를 제안한다. 이해 단계에서 입력 언어를 영어로 변환하고, 추론 단계에서 self-attention과 feed-forward 구조를 활용해 영어로 추론 및 다국어 지식을 결합한 후, 생성 단계에서 원래 언어로 출력을 생성한다. 이는 LLM이 다국어를 처리하는 내재적 메커니즘을 구조화한 것으로, 기존 연구와 달리 영어와 비영어 언어 간 상호작용을 명확히 설명한다. MWork는 PLND를 통해 라벨 없이 언어별 뉴런을 탐지하고, 이를 비활성화해 워크플로우의 각 단계를 검증할 수 있다.

기술적 접근법

주요 결과

의의 및 한계

MWork는 LLM의 다국어 처리 메커니즘을 체계적으로 설명하며, PLND는 라벨 없이 언어별 뉴런을 탐지하는 새로운 방법을 제시한다. 이는 다국어 능력을 특정 언어에 집중적으로 향상시키는 데 효과적이다. 그러나 본 연구는 특정 모델(예: Chinese Llama)에서는 MWork가 적용되지 않는 경우도 보고하고 있어, 모든 LLM에 일반화하기는 어려운 한계가 있다. 또한, 언어별 뉴런의 구체적인 분포나 영향은 모델에 따라 달라질 수 있으므로, 보다 다양한 모델에서의 검증이 필요하다.

실용적 활용

MWork와 PLND는 다국어 LLM의 특정 언어 성능을 정밀하게 향상시키는 데 활용할 수 있다. 특히, 저자원 언어 개선에 효과적이며, 번역, 요약, QA 등 다국어 NLP 태스크에서 활용 가능하다. 또한, 학습 데이터가 제한된 상황에서도 높은 성능 향상을 기대할 수 있어, 글로벌 기업의 다국어 서비스 개선에 유용하다.