跳到正文
arXiv:cs.LG· Katherine Van Koevering, Jon Kleinberg·· 5 小时前AI 评分37

LLM 的随机性研究:大语言模型生成随机序列时复现并放大人类偏差

Heads, Tails, and AI Fails: LLMs, Randomness, and Human Judgments

AI 导读

研究用模拟抛硬币范式测试当代 LLM 生成二元随机序列的能力,发现模型复现了过度交替、厌恶长连串、首掷偏差等人类随机性偏差,且往往将其放大或引入模型特有扭曲。提高 temperature 能减少部分僵化模式,但无法消除系统性结构;提示词框架、续写测试、语料检索与残差流探针显示,过度交替偏差无法仅用直接记忆或 tokenization 解释。

正文

View PDF HTML (experimental)

Abstract:Randomness is central to human cognition and to many applications in which large language models are deployed, yet probabilistic token generation does not imply that LLMs can produce unbiased random sequences. We study how contemporary LLMs generate binary random sequences using the classic behavioral-science paradigm of simulated coin flips. Across single flips, 20-flip sequences, n-gram statistics, run lengths, alternation rates, and next-flip predictability, we compare model outputs to both true Bernoulli baselines and human data from prior work. We find that LLMs reproduce several canonical human randomness biases, including over-alternation, aversion to long runs, and first-flip biases, but often amplify them or introduce model-specific distortions. Increasing temperature reduces some rigid patterns but does not eliminate systematic structure. We further investigate over-alternation through prompt-framing experiments, continuation tests, corpus searches, and residual-stream probes, finding evidence that this bias is not explained by direct memorization or tokenization alone. These results show that LLMs are unreliable generators of randomness and imperfect simulators of human randomness judgments, with implications for applications requiring stochasticity, fair sampling, or behavioral simulation.
Subjects: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
Cite as: arXiv:2406.00092 [cs.AI]
  (or arXiv:2406.00092v2 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2406.00092

arXiv-issued DOI via DataCite

Submission history

From: Katherine Van Koevering [view email]
[v1] Fri, 31 May 2024 17:56:07 UTC (1,278 KB)
[v2] Thu, 1 Oct 2026 18:08:19 UTC (742 KB)

来源:arXiv:cs.LG · arxiv.org