model-evaluation llms personalization user-profiles faithfulness-taxonomy stereotypical-profiles attribute-fabrication over-inference
Abstract
Personalized LLMs with persistent memory are increasingly deployed, yet the faithfulness of their user models remains unexamined. We study over-inference (OI): the phenomenon where LLMs fabricate user attributes beyond what evidence supports. We introduce MirageBench, comprising 150 personas balanced across stereotypical, counter-stereotypical, and neutral profiles, 6 personalization tasks spanning an ``imagination gradient'', a four-way faithfulness taxonomy operationalized by an independent judge (validated against a blind human annotator on 400 claims: Cohen's kappa = 0.863 four-class, kappa = 0.900 binary), and a leaderboard of 12 models across 7 families on 143616 judged claims. We find that over-inference is pervasive: every one of the 12 models over-infers 35%--49% of its claims (cross-model mean 41.6%; claim-weighted 41.8%), with no model in this evaluation escaping it. Most strikingly, we surface a Self-Monitoring Inversion: at the model-selection level, models' self-assessed OI is negatively rank-correlated with their judge-measured OI (rho = -0.60, p = 0.044; exploratory, wide bootstrap CI [-0.90, +0.06], n = 12). The models that report the least over-inference tend to be flagged as fabricating the most, so self-reported confidence is a misleading signal for comparing models, even though within a single model self-audit still ranks that model's own claims moderately well (AUROC 0.58--0.83). We further show that OI is task-dependent (27%--59%) and that, in a multi-turn pilot, inferred attributes accumulate approximately linearly with little revision. MirageBench positions external verification, rather than model self-report, as a more reliable foundation for trustworthy personalization.
한국어 요약
한 줄 요약
개인화된 대형 언어 모델(LLM)이 사용자 정보를 과도하게 추론하는 문제를 분석하고, 자기 평가가 신뢰도를 왜곡하는 현상을 밝혀낸 연구.
핵심 기여도
- MirageBench: 150개 사용자 프로필, 6개 개인화 작업, 4단계 신뢰도 분류, 143,616개 판단된 주장으로 구성된 첫 번째 과도한 추론(OI) 벤치마크.
- 12개 모델(7개 가족)에서 OI 비율 평균 41.6%로, 모든 모델이 35% 이상의 OI 발생.
- Self-Monitoring Inversion 발견: 모델의 자기 평가 OI는 외부 판단 OI와 음의 상관관계 (ρ = -0.60, p = 0.044).
- 다중 턴 실험에서 9개 모델이 5~15개의 새로운 추론 속성 추가, 0.4~5%만 수정.
핵심 아이디어
개인화된 LLM이 사용자 정보를 과도하게 추론하는 문제(OI)는 기존의 사실적 환각이나 사회적 편견과 구별된다. 이 연구는 OI를 측정하고 비교하기 위해 MirageBench라는 체계적인 평가 프레임워크를 제시한다. 이는 4단계 신뢰도 분류(Grounded / Reasonable / Stereotype / Fabricated)와 독립적 판단자(Judge)를 통해 외부 검증을 기반으로 한다. 연구는 특히, 모델이 자신이 얼마나 정확하게 추론하고 있다고 주장하는 자기 평가(Self-Monitoring)가 외부 판단과 반대되는 경향이 있음을 밝혀내어, 이는 시스템 설계에 중요한 시사점을 제공한다.
기술적 접근법
- **MirageBench 구성**: 150개 사용자 프로필(스테레오타입, 반스테레오타입, 중립), 6개 개인화 작업, 4단계 신뢰도 분류.
- **평가 단계**:
- **Probe**: 모델에게 "이 3가지 사실을 바탕으로 유추할 수 있는 모든 것을 나열하라"고 명령.
- **Task**: 6개 작업 수행 후, 모델 스스로 주장 분류 (4단계).
- **Judge**: 독립적 판단자(Claude-Opus-4-7)가 동일한 분류 기준으로 주장 평가.
- **Accum**: 8라운드의 상호작용에서 추론 속성의 축적 및 수정 추적.
- **신뢰도 검증**: 400개 주장에 대해 블라인드 인간 어노테이터와 비교 (Cohen’s κ = 0.863, 0.900).
주요 결과
- **모델별 OI 비율**: 12개 모델 모두 35%~49%의 OI 발생 (평균 41.6%, 주장 가중 평균 41.8%).
- **Self-Monitoring Inversion**: 자기 평가 OI가 낮은 모델일수록 외부 판단 OI가 높음 (ρ = -0.60, p = 0.044).
- **작업별 OI 차이**: 구체적 작업(27%)에서 상상적 작업(59%)으로 증가.
- **추론 축적**: 9개 모델에서 5~15개의 새로운 속성 추가, 수정률 0.4~5% (상위 5개 모델).
의의 및 한계
이 연구는 개인화 LLM의 신뢰도 문제를 체계적으로 평가할 수 있는 첫 번째 벤치마크(MirageBench)를 제시하며, 자기 평가가 신뢰도를 왜곡할 수 있음을 실증적으로 밝혔다. 이는 시스템 설계에서 외부 검증과 추적 기반의 접근이 필요함을 시사한다. 그러나 연구는 2개 사용자 프로필과 특정 메모리 프롬프트를 사용한 편입 실험에 기반하므로, 일반화 가능성에 한계가 있을 수 있다. 또한, 판단자(Judge)의 인간과의 일관성은 높지만, 모든 주장에 대한 완전한 인간 검증은 이루어지지 않았다.
실용적 활용
이 연구는 개인화 AI 시스템을 설계하거나 사용하는 모든 분야에 적용 가능하다. 특히, 사용자 프로필을 생성하는 챗봇, 메모리 기반 대화 시스템, 맞춤형 추천 엔진 등에서 외부 검증 기반의 신뢰성 확보가 필요하다. MirageBench는 이러한 시스템의 평가와 개선에 기초 자료로 활용될 수 있다.