arXiv:cs.CL· Tapan Parikh·· 4 小时前
一项针对 44 个语言模型的“单词普查”:答案选择趋同性研究
The One-Word Census: Answer-Choice Conformity Across 44 Language Models
AI 导读
一项研究用 96 个单轮提示词测试了 105 个语言模型在“说出一个词”类任务中的答案趋同现象,发现“serendipity”一词被 46% 的模型选中。在 96 个类别中有 28 个类别单一答案占比超过 80%,且 87 个主流实验室模型的集中度不低于整体水平。轻量后训练和人格微调模型分歧最大,重度后训练的助手模型最趋同;监督微调是走向趋同的最大一步。
正文
Abstract:When a language model must choose one answer from a large space of equally valid options, which answer does it choose, and how often is it the answer every other model chooses? Asked to "pick a word," 105 language models from more than twenty labs chose serendipity 46% of the time. We measure this convergence, and each model's share in it, with 96 single-turn prompts that each name a category with many valid one-word answers ("Name a tree."), asked eight times per model and scored by exact match, with no embeddings and no judge. A model's answer-choice surprisal is the average -log2 probability of its answers under the pooled answers of all other models. In 28 of 96 categories a single answer takes at least 80% of all answers. The concentration does not depend on small or persona-tuned models: the 87 major-lab models are at least as concentrated as the full field. Lightly post-trained and persona-tuned models are the most divergent; heavily post-trained assistants from the major labs are the most conformist. Models that avoid the modal answer mostly land on the same runner-up. Within the major providers' lineages, release order shows no panel-wide trend once model tier is controlled; GPT, Gemini, Grok and Qwen become more conformist across releases, and Claude's generation-5 releases reverse. On open checkpoints of three post-training pipelines, supervised fine-tuning is the largest step toward the field's answers. Against human category-production norms, the field is more concentrated than people in 18 of 20 shared categories. All prompts, transcripts, and code are public.
| Comments: | v3: 105 models, 96 prompts, 8 runs. Adds a major-lab vs small-model split, a spread-free score, training-stage checkpoints and a human-norms comparison. Data, code and transcripts: this http URL tag consensus-arxiv-v3) |
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computers and Society (cs.CY) |
| Cite as: | arXiv:2607.12796 [cs.CL] |
| (or arXiv:2607.12796v3 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2607.12796 arXiv-issued DOI via DataCite |
Submission history
From: Tapan Parikh [view email]
[v1]
Tue, 14 Jul 2026 14:12:05 UTC (293 KB)
[v2]
Sat, 25 Jul 2026 18:18:32 UTC (297 KB)
[v3]
Wed, 7 Oct 2026 20:52:36 UTC (123 KB)
来源:arXiv:cs.CL · arxiv.org