llm-evaluation multi-turn-conversation benchmark-design role-playing-agents system-comparison llm-simulators user-simulation personalized-rubrics
Abstract
Role-playing agents (RPAs) have become one of the most important consumer applications of large language models. Users engage in multi-turn conversations with RPAs for experiences such as emotional comfort, making reliable evaluation essential for measuring capability, comparing systems, and guiding further improvement. Existing benchmarks, however, typically require an RPA to continue a fixed dialogue history and then evaluate the continuation using a fixed rubric detached from the user. We identify and empirically demonstrate two limitations of this design. First, an RPA's output is shaped by the preceding dialogue history, preventing a scientifically grounded assessment of its role-playing ability in real multi-turn settings. Second, user experience varies substantially across individuals, and conventional fixed rubrics need not align with user satisfaction. We therefore introduce PALATE (Person-Aligned LLM-Simulated-User Assessment with Tailored Evaluation), a scalable RPA benchmark built on user simulators. PALATE is accompanied by a pool of 300 character profiles. Its main evaluation trains five per-user simulators and lets them engage candidate RPAs in free-form, multi-turn conversations over a pre-frozen panel of character profiles. Alongside a general quality rubric, we construct personalized rubrics to measure user satisfaction; on held-out annotated data, the personalized rubrics show higher agreement with human judgments than the general rubric. In the main evaluation of 16 candidates, PALATE separately characterizes generic turn quality, long-horizon session capability, and per-user experience on multi-turn trajectories co-constructed by each candidate. It thereby produces interpretable evaluations of specific user-RPA pairs rather than compressing systems into a single user-independent ranking.
한국어 요약
한 줄 요약
PALATE는 사용자 맞춤형 시뮬레이션을 통해 대화형 역할극 에이전트를 평가하는 새로운 벤치마크를 제시한다.
핵심 기여도
- 기존 평가 방식의 두 가지 한계를 실증적으로 밝힘: 고정된 대화 히스토리와 사용자와 분리된 평가 체계.
- PALATE는 300개의 캐릭터 프로필과 5개의 사용자 시뮬레이터를 기반으로 대화형 평가를 수행.
- 개인화된 평가 체계가 일반 평가 체계보다 인간 평가와 높은 일치도(명시되지 않음)를 보임.
- 16개 후보 에이전트를 대상으로 한 평가에서 사용자-에이전트 쌍별 평가를 가능하게 함.
핵심 아이디어
기존 역할극 에이전트 평가 방식은 고정된 대화 히스토리와 일반적인 평가 체계를 사용하여 실제 사용자 경험과 벗어나 있다. 이에 따라, PALATE는 사용자 프로필에 맞춘 시뮬레이터를 통해 역동적인 대화를 생성하고, 사용자별로 맞춤 평가 체계를 적용한다. 이는 사용자 경험의 다양성을 반영하며, 에이전트의 장기적 대화 능력과 사용자별 만족도를 구분해 평가할 수 있게 한다. 특히, `per-user simulators`와 `personalized rubrics`라는 모듈을 통해 사용자 맞춤 평가를 구현한다.
기술적 접근법
- 300개의 캐릭터 프로필을 기반으로 사용자 시뮬레이터를 생성.
- 각 시뮬레이터는 후보 에이전트와 자유형 다중 턴 대화를 수행.
- 평가 체계는 일반적인 품질 체크와 사용자별 맞춤 평가 체계로 구성.
- 시뮬레이션 대화는 고정된 히스토리가 아닌 동적으로 생성됨.
- 평가 대상은 16개의 역할극 에이전트.
주요 결과
- 16개 후보 에이전트에 대해 PALATE는 일반 턴 품질, 장기 대화 능력, 사용자별 경험을 별도로 평가.
- 개인화된 평가 체계는 일반 평가 체계보다 인간 평가와 더 높은 일치도를 보임 (명시되지 않음).
- 사용자-에이전트 쌍별 평가를 통해 단일 순위 평가 대신 해석 가능한 결과를 제공.
의의 및 한계
PALATE는 역할극 에이전트 평가에서 사용자 경험의 다양성을 반영함으로써 보다 실제에 가까운 평가를 가능하게 한다. 특히, 사용자 맞춤 평가 체계는 기존 평가 방식의 한계를 극복하고, 개별 사용자-에이전트 상호작용을 정확히 평가할 수 있다. 그러나 사용자 시뮬레이터의 정확도나 캐릭터 프로필의 범위가 제한적일 수 있으며, 인간 평가와의 정확한 일치도 수치는 명시되지 않았다.
실용적 활용
PALATE는 대화형 AI 제품 개발, 사용자 맞춤형 서비스 설계, 역할극 에이전트 비교 평가 등에 활용 가능하다. 특히, 사용자 경험 중심의 AI 개선을 추구하는 산업과 연구 분야에서 유용하게 사용될 수 있다.