The Personalization Mirage: How LLMs Fabricate User Profiles, and Why Self-Monitoring Misleads

Yushi Sun, Yanjie Zhang, Rui Sheng

arXiv:2608.04570 · 2026-08-06 공개 · arXiv · PDF

model-evaluation llms personalization user-profiles faithfulness-taxonomy stereotypical-profiles attribute-fabrication over-inference

Abstract

Personalized LLMs with persistent memory are increasingly deployed, yet the faithfulness of their user models remains unexamined. We study over-inference (OI): the phenomenon where LLMs fabricate user attributes beyond what evidence supports. We introduce MirageBench, comprising 150 personas balanced across stereotypical, counter-stereotypical, and neutral profiles, 6 personalization tasks spanning an ``imagination gradient'', a four-way faithfulness taxonomy operationalized by an independent judge (validated against a blind human annotator on 400 claims: Cohen's kappa = 0.863 four-class, kappa = 0.900 binary), and a leaderboard of 12 models across 7 families on 143616 judged claims. We find that over-inference is pervasive: every one of the 12 models over-infers 35%--49% of its claims (cross-model mean 41.6%; claim-weighted 41.8%), with no model in this evaluation escaping it. Most strikingly, we surface a Self-Monitoring Inversion: at the model-selection level, models' self-assessed OI is negatively rank-correlated with their judge-measured OI (rho = -0.60, p = 0.044; exploratory, wide bootstrap CI [-0.90, +0.06], n = 12). The models that report the least over-inference tend to be flagged as fabricating the most, so self-reported confidence is a misleading signal for comparing models, even though within a single model self-audit still ranks that model's own claims moderately well (AUROC 0.58--0.83). We further show that OI is task-dependent (27%--59%) and that, in a multi-turn pilot, inferred attributes accumulate approximately linearly with little revision. MirageBench positions external verification, rather than model self-report, as a more reliable foundation for trustworthy personalization.

한국어 요약

한 줄 요약

개인화된 대형 언어 모델(LLM)이 사용자 정보를 과도하게 추론하는 문제를 분석하고, 자기 평가가 신뢰도를 왜곡하는 현상을 밝혀낸 연구.

핵심 기여도

핵심 아이디어

개인화된 LLM이 사용자 정보를 과도하게 추론하는 문제(OI)는 기존의 사실적 환각이나 사회적 편견과 구별된다. 이 연구는 OI를 측정하고 비교하기 위해 MirageBench라는 체계적인 평가 프레임워크를 제시한다. 이는 4단계 신뢰도 분류(Grounded / Reasonable / Stereotype / Fabricated)와 독립적 판단자(Judge)를 통해 외부 검증을 기반으로 한다. 연구는 특히, 모델이 자신이 얼마나 정확하게 추론하고 있다고 주장하는 자기 평가(Self-Monitoring)가 외부 판단과 반대되는 경향이 있음을 밝혀내어, 이는 시스템 설계에 중요한 시사점을 제공한다.

기술적 접근법

주요 결과

의의 및 한계

이 연구는 개인화 LLM의 신뢰도 문제를 체계적으로 평가할 수 있는 첫 번째 벤치마크(MirageBench)를 제시하며, 자기 평가가 신뢰도를 왜곡할 수 있음을 실증적으로 밝혔다. 이는 시스템 설계에서 외부 검증과 추적 기반의 접근이 필요함을 시사한다. 그러나 연구는 2개 사용자 프로필과 특정 메모리 프롬프트를 사용한 편입 실험에 기반하므로, 일반화 가능성에 한계가 있을 수 있다. 또한, 판단자(Judge)의 인간과의 일관성은 높지만, 모든 주장에 대한 완전한 인간 검증은 이루어지지 않았다.

실용적 활용

이 연구는 개인화 AI 시스템을 설계하거나 사용하는 모든 분야에 적용 가능하다. 특히, 사용자 프로필을 생성하는 챗봇, 메모리 기반 대화 시스템, 맞춤형 추천 엔진 등에서 외부 검증 기반의 신뢰성 확보가 필요하다. MirageBench는 이러한 시스템의 평가와 개선에 기초 자료로 활용될 수 있다.