跳到正文
原文
arXiv:cs.AI(全量分类)· Fan Huang, Minsuk Kim, C. Tyler Diggans, Filippo Radicchi·· 5 小时前AI 评分40

评估 LLM 生成的偏好分布

Evaluating LLM-Generated Preference Distributions

AI 导读

研究系统分析了 LLM 在航空旅行、餐厅和消费品三个领域生成的偏好分布,发现九个开源权重模型均表现出自洽性,最可能结果在重复采样下快速稳定。但不同模型家族和规模之间存在显著不一致,即便最可能结果也鲜有共识,且该模式在温度变化、贪婪解码及提示词与顺序扰动下依然稳健。结果表明,结果更多受模型选择而非提示词措辞影响,挑战了能力足够强的 LLM 作为受访者替代品会产生相似偏好分布的假设。

正文

View PDF HTML (experimental)

Abstract:Large Language Models (LLMs) are increasingly used as probabilistic generators for simulation, synthetic data generation, and decision support in settings where real-world data are unavailable. Yet, the structure and reliability of the distributions they produce remain understudied. Here, we systematically analyze LLM-generated distributions of preferences for air travel, restaurants, and consumer products. Encouragingly, all models considered in our analysis exhibit self-coherence, with the most probable outcomes stabilizing rapidly under repeated sampling. At the same time, we observe substantial discordance across both model families and scales, with little consensus even among their most probable outcomes. These patterns hold across nine open-weight models, three choice domains, and show robustness under temperature changes, greedy decoding, and perturbations of prompt and ordering. Our findings indicate that outcomes are influenced more by the choice of model than by the wording of the prompt, challenging the common assumption that sufficiently capable LLMs produce similar preference distributions when used as stand-ins for survey respondents.
Subjects: Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.01000 [cs.AI]
  (or arXiv:2610.01000v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.01000

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Fan Huang [view email]
[v1] Thu, 1 Oct 2026 03:40:05 UTC (3,105 KB)

来源:arXiv:cs.AI(全量分类) · arxiv.org