跳到正文
arXiv:cs.CL· Liwei Jiang, Yuanjun Chai, Margaret Li, Mickel Liu, Raymond Fok, Nouha Dziri, Yulia Tsvetkov, Maarten Sap, Alon Albalak, Yejin Choi·· 10 小时前AI 评分63

Artificial Hivemind 论文发布 Infinity-Chat 数据集,揭示大语言模型开放性生成趋同现象

Artificial Hivemind: The Open-Ended Homogeneity of Language Models (and Beyond)

AI 导读

arXiv 论文(NeurIPS 2025 D&B Oral)提出 Infinity-Chat 数据集,包含 26K 条开放性真实用户查询和 31,250 条人工标注,并给出覆盖 6 大类、17 子类的开放性提示分类法。

正文

View PDF HTML (experimental)

Abstract:Language models (LMs) often struggle to generate diverse, human-like creative content, raising concerns about the long-term homogenization of human thought through repeated exposure to similar outputs. Yet scalable methods for evaluating LM output diversity remain limited, especially beyond narrow tasks such as random number or name generation, or beyond repeated sampling from a single model. We introduce Infinity-Chat, a large-scale dataset of 26K diverse, real-world, open-ended user queries that admit a wide range of plausible answers with no single ground truth. We introduce the first comprehensive taxonomy for characterizing the full spectrum of open-ended prompts posed to LMs, comprising 6 top-level categories (e.g., brainstorm & ideation) that further breaks down to 17 subcategories. Using Infinity-Chat, we present a large-scale study of mode collapse in LMs, revealing a pronounced Artificial Hivemind effect in open-ended generation of LMs, characterized by (1) intra-model repetition, where a single model consistently generates similar responses, and more so (2) inter-model homogeneity, where different models produce strikingly similar outputs. Infinity-Chat also includes 31,250 human annotations, across absolute ratings and pairwise preferences, with 25 independent human annotations per example. This enables studying collective and individual-specific human preferences in response to open-ended queries. Our findings show that LMs, reward models, and LM judges are less well calibrated to human ratings on model generations that elicit differing idiosyncratic annotator preferences, despite maintaining comparable overall quality. Overall, INFINITY-CHAT presents the first large-scale resource for systematically studying real-world open-ended queries to LMs, revealing critical insights to guide future research for mitigating long-term AI safety risks posed by the Artificial Hivemind.
Comments: NeurIPS 2025 D&B Paper (Oral); Camera-Ready Version
Subjects: Computation and Language (cs.CL)
Cite as: arXiv:2510.22954 [cs.CL]
  (or arXiv:2510.22954v2 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2510.22954

arXiv-issued DOI via DataCite

Submission history

From: Liwei Jiang [view email]
[v1] Mon, 27 Oct 2025 03:16:21 UTC (38,319 KB)
[v2] Tue, 6 Oct 2026 17:33:55 UTC (21,204 KB)

来源:arXiv:cs.CL · arxiv.org