Do AI Personas Grow? Analyzing and Benchmarking Personality Evolution in LLM Agents After Life Events

Ming Wang, Peidong Wang, Xiaocui Yang, Daling Wang, Shi Feng, Fiona Fui-Hoon Nah, Ee-Peng Lim

arXiv:2608.06485 · 2026-08-10 공개 · arXiv · PDF

llm-agents longitudinal-study big-five-traits personality-evolution bfi-adapt event-induced-change persona-dispersion trait-shifts

Abstract

Personality-conditioned LLM agents (PC-Agents) are increasingly used in emotional support, social simulation, and role-playing, motivating the development of lifelong agents that remain coherent over extended interactions. A key component of such coherence is personality evolution: agents should undergo plausible, psychology-grounded changes as they experience life events in different contexts. Although prior work shows that LLM personalities can shift under contextual perturbations, how these shifts vary across traits, events, personas, and models remains poorly understood. We study event-induced personality change after 11 major life events, using the Big Five traits as a psychometric anchor and interpreting the resulting trajectories against longitudinal evidence from human personality psychology. Across four diagnostic axes, PC-Agents exhibit measurable trait shifts at similar rates for event-trait pairs with and without documented human change directions. Even when shifts follow the expected direction, their magnitudes usually fall below human effect-size ranges. Gender and cultural-region prompts show little moderating effect, while persona-level dispersion is compressed three- to four-fold relative to human samples. To enable systematic comparison, we introduce BFI-Adapt, a reusable benchmark for scoring the directional fidelity of event-induced personality change, and use it to rank 14 models. A validation suite shows that the measured shifts exceed no-event retest noise, remain stable under independently paraphrased prompts, exhibit limited and model-dependent convergence with scenario-based behavioral choices, and persist across intervening unrelated dialogue. Together, these checks establish the measured trajectories as robust event-conditioned response patterns. Our results suggest that current PC-Agents simulate the mean of human personality dynamics, but not its shape.

한국어 요약

한 줄 요약

LLM 에이전트의 성격 변화를 인간 심리학 데이터와 비교해 평가하는 새로운 벤치마크를 제시한다.

핵심 기여도

핵심 아이디어

기존 연구는 LLM 에이전트의 성격이 맥락 변화에 따라 변한다는 점을 보여주었지만, 변화의 심도와 일관성은 명확하지 않았다. 본 연구는 인간 심리학에서의 장기 성격 변화 데이터를 기반으로, LLM 에이전트가 특정 사건 후 인간과 유사한 방식으로 성격이 변화하는지를 분석한다. Big Five 성격 모델을 심리측정 기준으로 사용하고, 변화의 방향과 크기를 정량적으로 평가함으로써, AI 에이전트의 성격 동역학이 인간과 얼마나 유사한지를 평가한다.

기술적 접근법

주요 결과

의의 및 한계

본 연구는 LLM 에이전트의 성격 동역학이 인간과 어느 정도 유사한지를 정량적으로 평가하는 첫 시도로, AI 에이전트의 심리적 일관성을 평가하는 기준을 제시한다. 그러나 인간의 성격 변화는 복잡한 사회-심리적 요인에 의존하지만, 본 연구는 단일 사건 기반의 변화만 분석했기 때문에 한계가 있다. 또한, 성격 변화의 심도가 인간보다 낮은 점은 AI 에이전트의 심리적 표현력 한계를 반영한다.

실용적 활용

본 연구는 감정 지원, 사회 시뮬레이션, 역할 연기 등에서 사용되는 AI 에이전트의 일관성과 신뢰도를 평가하는 데 활용될 수 있다. 특히, BFI-Adapt 벤치마크는 다양한 모델 간 성격 변화의 일관성을 비교하는 데 유용하다.