A Roadmap to Pluralistic Alignment

Taylor Sorensen, Jared Moore, Jillian R. Fisher, Mitchell Gordon, Niloofar Mireshghallah, Christopher Rytting, Andre Ye, Liwei Jiang, Ximing Lu, Nouha Dziri, Tim Althoff, Yejin Choi

arXiv:2402.05070 · 2026-07-27 공개 · arXiv · PDF

language-models overton-pluralism alignment-techniques pluralistic-alignment steerably-pluralistic distributionally-pluralistic multi-objective-benchmarks trade-off-steerable

Abstract

With increased power and prevalence of AI systems, it is ever more critical that AI systems are designed to serve all, i.e., people with diverse values and perspectives. However, aligning models to serve pluralistic human values remains an open research question. In this piece, we propose a roadmap to pluralistic alignment, specifically using language models as a test bed. We identify and formalize three possible ways to define and operationalize pluralism in AI systems: 1) Overton pluralistic models that present a spectrum of reasonable responses; 2) Steerably pluralistic models that can steer to reflect certain perspectives; and 3) Distributionally pluralistic models that are well-calibrated to a given population in distribution. We also formalize and discuss three possible classes of pluralistic benchmarks: 1) Multi-objective benchmarks, 2) Trade-off steerable benchmarks, which incentivize models to steer to arbitrary trade-offs, and 3) Jury-pluralistic benchmarks which explicitly model diverse human ratings. We use this framework to argue that current alignment techniques may be fundamentally limited for pluralistic AI; indeed, we highlight empirical evidence, both from our own experiments and from other work, that standard alignment procedures might reduce distributional pluralism in models, motivating the need for further research on pluralistic alignment.

한국어 요약

한 줄 요약

AI 시스템이 다양한 인간 가치를 반영하도록 정렬하는 "플러럴리즘 정렬"을 위한 구체적 로드맵을 제안한다.

핵심 기여도

핵심 아이디어

AI 정렬의 목표는 단일 가치에 맞추는 것이 아니라, 다양한 인간 가치를 반영하는 "플러럴리즘"을 달성하는 것이다. 이 논문은 이에 대한 구체적 정의와 평가 방법을 제시한다. Overton pluralism은 다양한 합리적 응답을 제공하는 모델을 의미하며, Steerably pluralism은 특정 관점으로 조정 가능한 모델을, Distributionally pluralism은 특정 인구 분포에 잘 맞춘 모델을 의미한다. 이는 단순히 "다양성"을 추구하는 것이 아니라, 각각의 상황에 따라 다른 정의가 필요하다는 통찰을 반영한다. 또한, 기존 정렬 기법이 Distributionally pluralism을 감소시킬 수 있다는 점에서, 새로운 정렬 방법론이 필요하다는 결론을 도출한다.

기술적 접근법

주요 결과

의의 및 한계

이 연구는 AI 정렬의 새로운 차원인 "플러럴리즘"을 명확히 정의하고, 이를 평가할 수 있는 벤치마크를 제안함으로써, 기존 정렬 기법의 한계를 드러낸다. 특히, 정렬 과정이 모델의 분포 다양성을 줄일 수 있다는 점은 중요한 학술적 통찰이다. 그러나, 이 연구는 실험 데이터가 제한적이며, 실제 대규모 모델에서의 적용 가능성은 추가 연구가 필요하다는 한계가 있다.

실용적 활용

이 연구는 사회적 가치가 다양한 환경에서 AI를 사용하는 산업(예: 교육, 정책, 미디어)에 적용 가능하다. 또한, 모델이 다양한 사용자 그룹의 의견을 반영하도록 설계해야 하는 연구 분야(예: 다문화 AI, 공정성 연구)에서도 활용 가능하다.