Political Compass or Spinning Arrow? Towards More Meaningful Evaluations for Values and Opinions in Large Language Models

Paul Röttger, Valentin Hofmann, Valentina Pyatkin, M. Hinck, Hannah Rose Kirk, Hinrich Schutze, Dirk Hovy

arXiv:2402.16786 · 2026-07-27 공개 · arXiv · PDF

llm-evaluation llm-bias multiple-choice political-compass-test values-ops survey-questions unconstrained-evaluation paraphrase-robustness

Abstract

Much recent work seeks to evaluate values and opinions in large language models (LLMs) using multiple-choice surveys and questionnaires. Most of this work is motivated by concerns around real-world LLM applications. For example, politically-biased LLMs may subtly influence society when they are used by millions of people. Such real-world concerns, however, stand in stark contrast to the artificiality of current evaluations: real users do not typically ask LLMs survey questions. Motivated by this discrepancy, we challenge the prevailing constrained evaluation paradigm for values and opinions in LLMs and explore more realistic unconstrained evaluations. As a case study, we focus on the popular Political Compass Test (PCT). In a systematic review, we find that most prior work using the PCT forces models to comply with the PCT's multiple-choice format. We show that models give substantively different answers when not forced; that answers change depending on how models are forced; and that answers lack paraphrase robustness. Then, we demonstrate that models give different answers yet again in a more realistic open-ended answer setting. We distill these findings into recommendations and open challenges in evaluating values and opinions in LLMs.

한국어 요약

한 줄 요약

LLM의 가치관 평가에서 강제 선택형 평가가 현실과 동떨어진 결과를 유발한다는 점을 밝히고, 보다 현실적인 평가 방법을 제안한다.

핵심 기여도

핵심 아이디어

기존 연구는 LLM의 가치관을 평가하기 위해 설문 조사 형식의 다중 선택 질문을 사용하지만, 실제 사용자 행동과는 거리가 있다. 이에 따라, 연구자들은 LLM이 답변을 강제로 선택하도록 유도하는 방식을 사용하지만, 이는 모델의 답변을 왜곡할 수 있다. 연구는 PCT를 사례로, 강제 평가가 결과의 불안정성과 일반화 불가능성을 초래함을 보여준다. 핵심 통찰은, LLM이 단일한 가치관을 가지는 것이 아니라, 입력 문맥에 따라 다양한 입장을 표현할 수 있다는 점이다. 이는 LLM이 인간처럼 단일 인격을 갖는 것이 아니라, 다양한 시뮬라크라의 중첩(superposition)을 표현한다는 관점과 일치한다.

기술적 접근법

주요 결과

의의 및 한계

이 연구는 LLM의 가치관 평가가 기존 방식으로는 신뢰할 수 없다는 점을 강조하며, 평가 형식이 결과에 큰 영향을 미친다는 점을 실증적으로 보여준다. 이는 LLM이 단일한 가치관을 갖는 것이 아니라, 다양한 시뮬라크라의 중첩을 표현한다는 개념과 일치하며, 평가 방법론의 재고를 촉구한다. 그러나 연구는 특정 평가 도구(PCT)에만 초점을 맞춘 한계가 있으며, 다른 평가 도구나 상황에서의 일반화 가능성은 명시되지 않았다. 또한, 개방형 평가에서도 불안정성이 남아 있다는 점에서, LLM이 어떤 방식으로든 일관된 가치관을 표현하는지는 여전히 논란이 있다.

실용적 활용

이 연구는 LLM이 실제 사용 환경에서 어떻게 가치관을 표현하는지를 이해하고, 이를 바탕으로 보다 현실적인 평가 시스템을 설계하는 데 기여할 수 있다. 특히, 사용자와의 대화에서 가치관이 어떻게 드러나는지를 평가하는 데 유용하며, 정책 결정, 교육, 미디어 등 다양한 분야에서 LLM의 신뢰성과 편향성을 평가하는 데 활용될 수 있다.