Investigating Cultural Alignment of Large Language Models

Badr AlKhamissi, Muhammad N. ElNokrashy, Mai Alkhamissi, Mona T. Diab

arXiv:2402.13231 · 2026-07-27 공개 · arXiv · PDF

llm-evaluation large-language-models cross-lingual-transfer survey-simulation cultural-alignment language-culture-relationship multilingual-pretraining anthropological-prompting

Abstract

The intricate relationship between language and culture has long been a subject of exploration within the realm of linguistic anthropology. Large Language Models (LLMs), promoted as repositories of collective human knowledge, raise a pivotal question: do these models genuinely encapsulate the diverse knowledge adopted by different cultures? Our study reveals that these models demonstrate greater cultural alignment along two dimensions -- firstly, when prompted with the dominant language of a specific culture, and secondly, when pretrained with a refined mixture of languages employed by that culture. We quantify cultural alignment by simulating sociological surveys, comparing model responses to those of actual survey participants as references. Specifically, we replicate a survey conducted in various regions of Egypt and the United States through prompting LLMs with different pretraining data mixtures in both Arabic and English with the personas of the real respondents and the survey questions. Further analysis reveals that misalignment becomes more pronounced for underrepresented personas and for culturally sensitive topics, such as those probing social values. Finally, we introduce Anthropological Prompting, a novel method leveraging anthropological reasoning to enhance cultural alignment. Our study emphasizes the necessity for a more balanced multilingual pretraining dataset to better represent the diversity of human experience and the plurality of different cultures with many implications on the topic of cross-lingual transfer.

한국어 요약

한 줄 요약

LLM의 문화적 정합성을 평가하고 언어와 사전학습 데이터의 영향을 분석하여, 문화적 다변성을 반영한 새로운 프롬프팅 방법을 제안한다.

핵심 기여도

핵심 아이디어

LLM은 단순히 언어를 학습하는 것이 아니라, 언어에 내재된 문화적 지식을 반영한다는 가정 하에 연구가 진행되었다. 이 연구는 문화적 정합성을 **사회학적 설문 조사 재현**을 통해 측정하며, 모델 응답과 실제 설문 응답 간 유사도를 비교하는 방식을 사용했다.

핵심 통찰은 **프롬프트 언어**와 **사전학습 데이터의 언어 구성**이 문화적 정합성에 큰 영향을 미친다는 점이다. 예를 들어, **아랍어**로 프롬프트를 주면 **아랍권 문화**에 더 가까운 응답이 나온다. 또한, **사전학습 데이터에 다양한 언어가 포함될수록** 문화적 다변성을 더 잘 반영하는 것으로 나타났다.

기술적 접근법

주요 결과

의의 및 한계

이 연구는 LLM이 단순히 언어를 학습하는 것이 아니라, **문화적 맥락**을 반영한다는 점을 강조하며, **다국어 사전학습 데이터**의 중요성을 입증한다. 특히, **Anthropological Prompting**은 문화적 맥락을 고려한 추론을 유도하는 새로운 접근법으로, 문화 정합성 향상에 기여할 수 있다.

그러나 한계도 존재한다. 예를 들어, **이집트 내 방언 차이**나 **미표현 인구의 다양성**을 완전히 반영하지 못했으며, **Modern Standard Arabic** 대신 **이집트 방언**을 사용하면 더 높은 정합성이 기대된다. 또한, **문화적 지식의 검증이 어려운 점**도 한계로 지적된다.

실용적 활용

이 연구는 **다문화 사회에서 사용되는 AI 시스템**의 설계에 중요한 시사점을 제공한다. 예를 들어, **글로벌 기업의 마케팅 전략**이나 **정부 정책 설계**에서 LLM이 특정 문화에 더 잘 맞는 응답을 제공할 수 있도록 도울 수 있다. 또한, **교육용 AI**나 **번역 시스템** 개발에도 문화적 정합성을 고려한 설계가 필요하다.