Do Synthetic Personas Predict Real Audience Response? A Sim-to-Real Study Where a No-Persona Baseline Beats Persona-Based Copy Simulation

Alexandre Cristovão Maiorano

arXiv:2609.25010 · 2026-09-23 공개 · arXiv · PDF

llm zero-shot sim-to-real gpt-4 persona-simulation gemini-models click-through-rate predictive-validity

Abstract

Marketers increasingly use large language models (LLMs) as "synthetic personas" to predict how an audience will react to a piece of copy before it ships, encouraged by evidence that profile-conditioned LLMs mimic human samples. But is that prediction actually valid against real behaviour - and does the persona machinery help? We present a sim-to-real validity study using the Upworthy Research Archive - thousands of headline A/B tests on shared real traffic, with measured click-through - as held-out ground truth. We compare a ten-persona panel, grounded in the real audience's demographics, against a no-persona zero-shot baseline that simply asks the model how likely a typical reader is to click. Two findings stand out. First, ground-truth reliability is the binding constraint: most A/B tests have no statistically distinguishable winner, so validity can only be measured on the reliable subset (n = 399). Second, and counter to the persona-simulation premise, persona conditioning degrades predictive validity: the no-persona baseline ranks variants markedly better (Kendall {\tau} = 0.361, a medium effect; top-1 accuracy 49.2%) than the persona panel ({\tau} = 0.084; top-1 34.6%), with non-overlapping confidence intervals. Asking the model directly taps an accurate population-level prior; forcing it to role-play specific personas injects bias and noise. The result replicates across three independent Upworthy splits, holds in direction on a different-domain news dataset, and is robust to seed, prompt phrasing, and model choice - across three Gemini tiers and a different model family (OpenAI gpt-4.1, significant paired gap). The takeaway: for predicting aggregate engagement, a plain LLM ranker beats persona simulation - synthetic personas are not merely a weak predictor, they are worse than not using them. All numbers regenerate from a public, artifact-first replication package.

한국어 요약

한 줄 요약

Upworthy A/B 테스트 데이터를 기반으로, 인공 인물(persona)을 사용한 마케팅 콘텐츠 시뮬레이션은 예측 정확도가 낮아, 단순 LLM 기반 순위 매기기가 더 효과적임을 밝힘.

핵심 기여도

핵심 아이디어

마케터들이 LLM을 "인공 인물"로 사용해 실제 대상자의 반응을 시뮬레이션하는 방식은 인기 있지만, 실제 행동 예측에 얼마나 유효한지는 검증되지 않았다. 본 연구는 Upworthy의 수천 개 A/B 테스트 데이터를 기반으로, 인공 인물 시뮬레이션이 실제 클릭률 예측에 도움이 되는지 평가했다. 핵심 통찰은, 인공 인물 조건을 주입하는 것이 오히려 모델의 정확도를 낮춘다는 점이다. 모델이 직접 인구 통계적 패턴을 학습한 상태에서 질문받는 것이, 특정 인물의 역할극을 강요하는 것보다 더 정확한 예측을 가능하게 한다는 것이다. 이는 LLM이 "인구 수준의 사전 정보(prior)"를 보유하고 있다는 사실을 반영한다.

기술적 접근법

주요 결과

의의 및 한계

본 연구는 마케팅 분야에서 LLM을 인공 인물로 사용하는 실용적 접근법의 한계를 명확히 드러냈다. 인공 인물은 모델의 정확도를 낮추며, 단순히 모델에 "평균 독자의 클릭 가능성"을 묻는 것이 더 효과적임을 보여준다. 이는 LLM이 인구 수준의 클릭 패턴을 학습하고 있다는 점을 시사한다. 그러나 인공 인물이 유용할 수 있는 상황(예: 시장 세분화, 질적 분석)은 여전히 존재하며, 본 연구는 순위 예측에만 한정된 결론이다. 한계점으로는, 인공 인물이 아닌 다른 유형의 조건(예: 감정, 문화적 배경)이 예측 성능에 어떤 영향을 미치는지에 대한 분석이 부재하다는 점을 들 수 있다.

실용적 활용

본 연구는 마케팅 콘텐츠 최적화 과정에서 인공 인물 시뮬레이션 대신, 단순 LLM 기반 순위 매기기를 사용할 것을 권장한다. 특히, 대규모 인구 집단의 반응을 예측하는 데에는 인공 인물이 불필요한 복잡성과 오류를 유발할 수 있다. 이는 광고, 뉴스, SNS 콘텐츠 등 다양한 마케팅 도메인에서 즉각적으로 적용 가능한 실용적 인사이트이다.