Beyond Borrowed Histories: Person-Aligned User Simulation for Interactive Role-Playing Evaluation

Yuhang Zhu, Mingxuan Du, Benfeng Xu, Jie Gao, Lingyun Yu, Hongtao Xie

arXiv:2607.27816 · 2026-08-01 공개 · arXiv · PDF

llm-evaluation multi-turn-conversation benchmark-design role-playing-agents system-comparison llm-simulators user-simulation personalized-rubrics

Abstract

Role-playing agents (RPAs) have become one of the most important consumer applications of large language models. Users engage in multi-turn conversations with RPAs for experiences such as emotional comfort, making reliable evaluation essential for measuring capability, comparing systems, and guiding further improvement. Existing benchmarks, however, typically require an RPA to continue a fixed dialogue history and then evaluate the continuation using a fixed rubric detached from the user. We identify and empirically demonstrate two limitations of this design. First, an RPA's output is shaped by the preceding dialogue history, preventing a scientifically grounded assessment of its role-playing ability in real multi-turn settings. Second, user experience varies substantially across individuals, and conventional fixed rubrics need not align with user satisfaction. We therefore introduce PALATE (Person-Aligned LLM-Simulated-User Assessment with Tailored Evaluation), a scalable RPA benchmark built on user simulators. PALATE is accompanied by a pool of 300 character profiles. Its main evaluation trains five per-user simulators and lets them engage candidate RPAs in free-form, multi-turn conversations over a pre-frozen panel of character profiles. Alongside a general quality rubric, we construct personalized rubrics to measure user satisfaction; on held-out annotated data, the personalized rubrics show higher agreement with human judgments than the general rubric. In the main evaluation of 16 candidates, PALATE separately characterizes generic turn quality, long-horizon session capability, and per-user experience on multi-turn trajectories co-constructed by each candidate. It thereby produces interpretable evaluations of specific user-RPA pairs rather than compressing systems into a single user-independent ranking.

한국어 요약

한 줄 요약

PALATE는 사용자 맞춤형 시뮬레이션을 통해 대화형 역할극 에이전트를 평가하는 새로운 벤치마크를 제시한다.

핵심 기여도

핵심 아이디어

기존 역할극 에이전트 평가 방식은 고정된 대화 히스토리와 일반적인 평가 체계를 사용하여 실제 사용자 경험과 벗어나 있다. 이에 따라, PALATE는 사용자 프로필에 맞춘 시뮬레이터를 통해 역동적인 대화를 생성하고, 사용자별로 맞춤 평가 체계를 적용한다. 이는 사용자 경험의 다양성을 반영하며, 에이전트의 장기적 대화 능력과 사용자별 만족도를 구분해 평가할 수 있게 한다. 특히, `per-user simulators`와 `personalized rubrics`라는 모듈을 통해 사용자 맞춤 평가를 구현한다.

기술적 접근법

주요 결과

의의 및 한계

PALATE는 역할극 에이전트 평가에서 사용자 경험의 다양성을 반영함으로써 보다 실제에 가까운 평가를 가능하게 한다. 특히, 사용자 맞춤 평가 체계는 기존 평가 방식의 한계를 극복하고, 개별 사용자-에이전트 상호작용을 정확히 평가할 수 있다. 그러나 사용자 시뮬레이터의 정확도나 캐릭터 프로필의 범위가 제한적일 수 있으며, 인간 평가와의 정확한 일치도 수치는 명시되지 않았다.

실용적 활용

PALATE는 대화형 AI 제품 개발, 사용자 맞춤형 서비스 설계, 역할극 에이전트 비교 평가 등에 활용 가능하다. 특히, 사용자 경험 중심의 AI 개선을 추구하는 산업과 연구 분야에서 유용하게 사용될 수 있다.