Large-Scale ChatBot Validation Through Customer Digital Twin Simulations

Cristovao Iglesias, Devesh Batra, Alankar Atreya, Stefan Wagner, Robert Hankache, Patrick Sinclair, Giulio Pelosio, Michael McMillan, Greig A. Cowan, Raad Khraishi

arXiv:2607.26060 · 2026-07-30 공개 · arXiv · PDF

llm-as-a-judge semantic-alignment digital-twins regulatory-compliance customer-service llm-validation adversarial-probing banking-applications

Abstract

LLM-based chatbots are transforming customer service in regulated domains such as banking, but scalable and cost-effective validation remains a critical barrier to safe deployment. We present a two-part contribution for large-scale chatbot validation. First, we introduce a methodology for creating high-fidelity synthetic customer agents (SCAs) as digital twins, grounded in real transactional and conversational data, that enables automatic generation and behavioral conditioning to simulate diverse customer profiles and interaction styles. Evaluation demonstrates that SCAs achieve high semantic alignment with real customers, low hallucination rates, and successful personality trait reproduction with controllable interventions. Second, we develop an SCA-based validation framework combining automated LLM-as-a-Judge evaluation, human expert testing, and adversarial probing. Scenario-based validation across emotional states, demographic groups, and linguistic factors confirms robust performance. Our approach was used to validate a customer facing chatbot at a leading UK bank, providing financial institutions with a scalable pathway toward regulatory compliance.

한국어 요약

한 줄 요약

고품질 디지털 트윈 기반 고객 시뮬레이션으로 은행용 챗봇의 대규모 검증이 가능해진다.

핵심 기여도

핵심 아이디어

기존 챗봇 검증 방식은 소규모 인간 실험에 의존해 비용이 높고 확장성이 낮아, 금융 등 규제 도메인에서의 안전한 배포를 어렵게 한다. 이를 해결하기 위해, 실제 거래 및 대화 데이터를 기반으로 한 고정밀 합성 고객 에이전트(SCA)를 디지털 트윈으로 생성하는 방법을 제안한다. SCA는 자동 생성과 행동 조건 설정을 통해 다양한 고객 프로필과 대화 스타일을 시뮬레이션하며, GPT-4.1 기반 모델을 사용해 의미 일치도가 높고 홀루시네이션이 낮은 대화를 생성한다.

또한, SCA를 기반으로 한 검증 프레임워크는 LLM-as-a-Judge 평가, 전문가 테스트, 적대적 탐색을 결합해 챗봇의 성능을 다각도로 평가한다. 특히, Big Five 성격 특성(IPIP-NEO-300)을 기반으로 SCA의 성격 조절이 가능하며, 감정 상태(예: 화난 고객)를 조정해 실제와 유사한 반응을 유도할 수 있다. 이는 과거 데이터를 기반으로 한 컨트라팩추얼 시나리오 생성을 가능하게 하며, 규제 준수와 공정성을 동시에 달성하는 데 기여한다.

기술적 접근법

주요 결과

의의 및 한계

이 연구는 규제 도메인에서 대규모 챗봇 검증을 위한 확장 가능한 프레임워크를 제시하며, 실제 은행 챗봇 검증에 적용되어 규제 준수를 촉진한다. 특히, SCA는 실제 고객 데이터를 기반으로 하므로, 기존 시뮬레이션 방법보다 더 높은 신뢰도를 제공한다. 또한, LLM-as-a-Judge와 적대적 탐색을 결합한 검증 방식은 챗봇의 공정성과 안정성을 체계적으로 평가할 수 있는 기반을 마련한다.

그러나, SCA는 실제 고객의 감정 표현(예: 신경증)을 완벽히 재현하지 못하는 한계가 있으며, 일부 오류는 LLM-as-Judge 평가를 통해 보완해야 한다. 또한, 특정 은행의 챗봇 검증 사례만 제시되어 다른 도메인으로의 확장 가능성은 추가 연구가 필요하다.

실용적 활용

이 연구는 금융, 보험, 의료 등 규제가 엄격한 도메인에서 챗봇의 대규모 검증을 가능하게 하며, 특히 고객 대응 챗봇의 공정성, 안정성, 규제 준수를 보장하는 데 유용하다. 또한, SCA 기반 시뮬레이션은 고객 행동 패턴 분석, 서비스 개선, 리스크 관리 등 다양한 실무 상황에서 활용 가능하다.