MatrAIx: Simulating the World with 8.3 Billion Persona Agents

Xiaomin Li, Yuexing Hao, Jianheng Hou, Jintao Huang, Qianfeng Wen, Shirley Huang, Yifan Liu, Xiaoyi Liu, Yilan Fan, Yijun Wang, Koutian Wu, Ruoqi Gao, Muhammad Ahmed Mohsin, Jing Tang, Brihi Joshi, Heming Liu, Zheyuan Deng, Zonglin Di, Sankalp Jajee, Jiuyao Lu, Zhiwei Zhang, Saksham Kapoor, Ishan Gupta, Yunhan Zhao, Chanwoo Park, Yucheng Lu, Bing Hu, Weihang Xiao, Aravind Mohan, Hanwen Xing, Runyu Zhang, Mihir Kulshreshtha, Yuanda Xu, Qianyu Zhu, Dianzhuo Wang, Yuxin Xiao, Bowen Jiang, Yongye Su, Wenhao Chai, Zuxin Liu, Lawrence Yunliang Chen, Xuandong Zhao, Ethan Ye, Shivam Patel, Jason Xie, Alex Martin Richmond, Weixiang Ding, Emre Okcular, Diya Mathew, Ziheng Wang, Rana M. Shahroz Khan, Zhejian Peng, Fang Wu, Fan Nie, Xinyang Han, Yubin Kim, Jiawei Zhang, Zhenting Qi, Huangyuan Su, Xu Pan, Abinitha Gourabathina, Hyewon Jeong, Hemanth Neelgund Ramesh, Kumail Alhamoud, Kimia Hamidieh, Zidi Xiong, Samuel Schmidgall, Pengrui Han, Yepeng Huang, Yongheng Wang, Bowen Yang, Alex Gu, Yuchu Wang, Akshay Paruchuri, Brenna Li, Hejie Cui, Jiayuan Ding, Chaosheng Dong, Jiahao Wang, Yixuan He, Chi Wang, Pamela Bhattacharya, Tianyi Peng, Paul Pu Liang, Mitchell Gordon, Yilun Du, Marinka Zitnik, James Zou, Prasanna Tambe, Philip Torr, Emily Fox, Asu Ozdaglar, Dawn Song

arXiv:2608.04205 · 2026-08-11 공개 · arXiv · PDF

llm-evaluation chatbot-evaluation persona-simulation user-diversity application-tasks persona-8b playground-framework feedback-analysis

Abstract

Human evaluation of AI systems and digital products is costly, slow, and difficult to scale. Offline evaluations are more scalable but often abstract away human diversity and interactive behavior. We therefore introduce MatrAIx, a population-scale simulated-user evaluation infrastructure for testing AI systems and digital products with heterogeneous users. MatrAIx has three core components: First, Persona 8B contains 8.3 billion persona records represented by 1,290 categorical dimensions. Records are either sampled from a dependency graph that preserves correlated attributes or derived from human-authored profiles. We release a quality-filtered coreset of approximately 1 million personas, comprising 599,847 human-grounded and 400,000 synthetic records. Second, the MatrAIx Playground provides four environments in which diverse users evaluate and interact with digital products: Survey, AI Chatbot, Web, and App. Third, MatrAIx provides 1,010 application tasks spanning more than 25 domains, including Commerce, Software, Finance, and Healthcare. We conducted 18,189 evaluation trials across eight representative tasks. Persona agents were powered by three LLMs: Claude Opus 4.8, GPT 5.5, and Claude Haiku 4.5. The resulting feedback captures how decisions and preferences vary across persona backgrounds, including hesitation after a price increase, willingness to continue after an AI assistant fails, and latency tolerance. We conducted two main validation studies: First, a 400-trial controlled study evaluated persona adherence across ten behavioral attributes and all four environments. The declared behavior was expressed or correctly suppressed in 366 trials (91.5%). Second, human and LLM judges evaluated the extraction quality of human-grounded personas. Overall, MatrAIx provides an end-to-end infrastructure for evaluating AI systems and digital products with diverse simulated human users.

한국어 요약

한 줄 요약

MatrAIx는 83억 개의 인물 기반 시뮬레이션을 통해 AI 시스템과 디지털 제품을 다각도로 평가하는 인프라를 제시한다.

핵심 기여도

핵심 아이디어

기존 오프라인 평가가 사용자 다양성과 상호작용을 무시하는 문제를 해결하기 위해, MatrAIx는 **Persona 8B**라는 대규모 인물 데이터셋을 기반으로 실제 사용자와 유사한 시뮬레이션을 구축한다. 이는 **dependency graph**를 활용한 합성 인물 생성과 **Wikipedia, Amazon Reviews, GSS** 등 6개 출처에서 추출한 실제 인물 기반 데이터를 결합한 방식이다. 인물 기록은 1,290개의 범주적 차원으로 구성되어 있으며, 이는 인물의 배경, 심리, 능력, 행동, 생활 방식 등을 포괄한다. MatrAIx Playground는 Survey, AI Chatbot, Web, App 4가지 환경에서 사용자와 시스템 간 상호작용을 시뮬레이션하며, 이는 다양한 디지털 제품과 AI 시스템의 평가에 활용된다.

기술적 접근법

주요 결과

의의 및 한계

MatrAIx는 AI 시스템과 디지털 제품의 사용자 중심 평가를 대규모로 가능하게 하며, **사전 배포 검증**, **서브그룹 분석**, **스트레스 테스트**, **버전 비교**에 활용 가능하다. 특히, **Persona 8B**는 실제 인물 데이터와 합성 데이터를 결합한 방식으로, 사용자 다양성을 효과적으로 반영한다. 그러나 **인물 에이전트 모델 간 결과 차이**가 있을 수 있으며, **실제 사용자와의 차이**는 여전히 존재하므로, 중요한 결론은 **인간 실험으로 검증**되어야 한다. 또한, **모델의 제한된 행동 표현**이나 **복잡한 상황 대응 능력**은 한계로 작용할 수 있다.

실용적 활용

MatrAIx는 소비자 행동 분석, AI 챗봇 성능 평가, 웹/앱 사용자 경험 테스트, 가격 민감도 조사 등 다양한 산업 분야에서 활용 가능하다. 특히, **AI 제품 개발 초기 단계에서 사용자 반응을 시뮬레이션**하여 문제를 사전에 탐지하는 데 유용하며, **다양한 사용자 그룹에 대한 분석**을 통해 맞춤형 개선 방안을 도출할 수 있다.