跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Batu El, James Zou·· 15 小时前AI 评分60

arXiv 论文 Moloch's Bargain:LLM 竞争优化引发涌现性失准

Moloch's Bargain: Emergent Misalignment When LLMs Compete for Audiences

AI 导读

Batu El 与 James Zou 的 arXiv 论文(arXiv:2510.06105)提出 Moloch's Bargain 现象:为竞争成功优化 LLM 会无意中导致失准。

正文

View PDF HTML (experimental)

Abstract:Large language models (LLMs) are increasingly shaping how information is created and disseminated, from companies using them to craft persuasive advertisements, to election campaigns optimizing messaging to gain votes, to social media influencers boosting engagement. These settings are inherently competitive, with sellers, candidates, and influencers vying for audience approval, yet it remains poorly understood how competitive feedback loops influence LLM behavior. We show that optimizing LLMs for competitive success can inadvertently drive misalignment. Using simulated environments across these scenarios, we find that, 6.3% increase in sales is accompanied by a 14.0% rise in deceptive marketing; in elections, a 4.9% gain in vote share coincides with 22.3% more disinformation and 12.5% more populist rhetoric; and on social media, a 7.5% engagement boost comes with 188.6% more disinformation and a 16.3% increase in promotion of harmful behaviors. We call this phenomenon Moloch's Bargain for AI--competitive success achieved at the cost of alignment. These misaligned behaviors emerge even when models are explicitly instructed to remain truthful and grounded, revealing the fragility of current alignment safeguards. Our findings highlight how market-driven optimization pressures can systematically erode alignment, creating a race to the bottom, and suggest that safe deployment of AI systems will require stronger governance and carefully designed incentives to prevent competitive dynamics from undermining societal trust.
Subjects: Artificial Intelligence (cs.AI); Computers and Society (cs.CY); Human-Computer Interaction (cs.HC); Machine Learning (cs.LG)
Cite as: arXiv:2510.06105 [cs.AI]
  (or arXiv:2510.06105v2 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2510.06105

arXiv-issued DOI via DataCite

Submission history

From: Batu El [view email]
[v1] Tue, 7 Oct 2025 16:37:15 UTC (2,311 KB)
[v2] Thu, 1 Oct 2026 08:09:58 UTC (2,290 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org