Towards Measuring and Modeling “Culture” in LLMs: A Survey

Farid Adilazuarda, Sagnik Mukherjee, Pradhyumna Lavania, Siddhant Singh, Ashutosh Dwivedi, Alham Fikri Aji, Jacki O'Neill, Ashutosh Modi, M. Choudhury

arXiv:2403.15412 · 2026-07-27 공개 · arXiv · PDF

llm model-evaluation survey llm-applications bias-measurement cultural-representation probing-methods semantic-proxies

Abstract

We present a survey of more than 90 recent papers that aim to study cultural representation and inclusion in large language models (LLMs). We observe that none of the studies explicitly define “culture, which is a complex, multifaceted concept; instead, they probe the models on some specially designed datasets which represent certain aspects of “culture”. We call these aspects the proxies of culture, and organize them across two dimensions of demographic and semantic proxies. We also categorize the probing methods employed. Our analysis indicates that only certain aspects of “culture,” such as values and objectives, have been studied, leaving several other interesting and important facets, especially the multitude of semantic domains (Thompson et al., 2020) and aboutness (Hershcovich et al., 2022), unexplored. Two other crucial gaps are the lack of robustness of probing techniques and situated studies on the impact of cultural mis- and under-representation in LLM-based applications.

한국어 요약

한 줄 요약

90개 이상의 논문을 바탕으로 LLM에서 문화 표현과 포용성을 조사하고, 문화의 정의 부재와 탐색 방법의 한계를 지적한다.

핵심 기여도

핵심 아이디어

LLM에서 문화를 탐색하기 위해 연구자들은 **문화 프록시**라는 개념을 도입하여, 특정 데이터셋을 통해 문화적 특성을 추적하고 있다. 이는 문화라는 복합적 개념을 단일 정의 없이 다차원적으로 접근하려는 시도이다. 그러나 문화는 **맥락적이고 상황적**(situated)이며, 디지털 텍스트는 이를 **두터운 묘사**(thick description)로 포착하기 어렵다는 점에서 한계가 있다. 특히, **미디지털화된 문화**(digitally under-represented cultures)는 외부자의 **얇은 묘사**(thin description)로 표현되어 편향과 고정관념을 심화시킬 수 있다. 이는 LLM이 문화적 미묘함을 정확히 이해하는 데 장애가 된다.

기술적 접근법

주요 결과

의의 및 한계

이 연구는 LLM에서 문화 표현의 중요성을 강조하고, 문화 탐색 방법의 체계적 분류를 제시함으로써 향후 연구의 기초를 제공한다. 특히, **문화의 정의 부재**와 **탐색 기법의 로버스트성 부족**은 학술적 발전의 핵심 과제이다. 그러나 본 연구는 **정량적 수치**나 **비교 실험 결과**를 제시하지 않아, 실제 모델 성능 개선 여부를 평가하기 어렵다는 한계가 있다.

실용적 활용

이 연구는 LLM이 다양한 문화적 맥락에서 **신뢰성 있게 작동**하도록 설계하는 데 기여할 수 있다. 예를 들어, **글로벌 고객 지원 시스템**이나 **문화적으로 민감한 콘텐츠 생성**에 활용될 수 있으며, **문화 다양성 보존**과 **편향 감소**를 위한 모델 개선 방향을 제시한다.