arXiv:cs.AI· Dongryeol Lee, Weronika {\L}ajewska, Leonardo Perelli, Saab Mansour·· 6 小时前AI 评分32
用数据驱动人格模拟问卷调查:数据访问机制下模拟对齐的实证研究
Data-Driven Personas for Survey Simulation: Insights into Simulation Alignment Across Data-Access Regimes
AI 导读
研究提出从异构匿名公共行为数据中诱导人格(persona),用于让 LLM 智能体模拟特定人口群体的问卷回答。结果发现,来自域外数据的人格很少优于仅依赖基本人口信息的模拟,主要源于人群不匹配;但当人格被准确分配到目标人口群体时,对齐度显著提升。来自目标域问卷数据的人格随问题历史增多而泛化更好,说明更丰富的行为证据能带来更稳定的人格特质推断。
正文
Abstract:Large language models (LLMs) offer new opportunities for public opinion research by enabling early prediction of survey responses, potentially reducing the cost and time of traditional surveys. However, many existing steering approaches rely on target-domain human data for fine-tuning or prompting that is costly to collect and raises privacy concerns. In this paper, we study demographic group-level survey simulation, where personas induced from heterogeneous, anonymized public behavioral data condition agents that simulate responses of individuals from specific demographic groups. We examine whether representative personas can be induced from diverse sources and analyze how the domain, scale, and granularity of the source data affect survey simulation alignment. We find that personas induced from out-of-domain sources rarely outperform simulations conditioned only on basic demographic information, largely due to population mismatch. However, when personas are accurately assigned to the target demographic groups, alignment improves substantially. Finally, personas induced from target-domain survey data generalize better as more survey question history becomes available, suggesting that richer behavioral evidence enables more stable persona trait inference that transfers to better unseen questions simulation alignment.
| Comments: | Work accepted at REALM EMNLP 2026 |
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.05828 [cs.AI] |
| (or arXiv:2610.05828v2 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.05828 arXiv-issued DOI via DataCite |
Submission history
From: Dongryeol Lee [view email]
[v1]
Mon, 5 Oct 2026 05:30:27 UTC (5,992 KB)
[v2]
Tue, 6 Oct 2026 07:14:13 UTC (5,992 KB)
来源:arXiv:cs.AI · arxiv.org