跳到正文
arXiv:cs.LG· Jessica Dierking, Itai Shapira, Niclas Boehmer·· 3 小时前AI 评分45

后训练中丢失了什么?默认坍缩与多元视角下上下文可引导性的丧失

What Is Lost in Post-Training? Default Collapse and the Loss of In-Context Steerability Across Diverse Perspectives

AI 导读

研究发现,后训练不仅收窄大语言模型表达的立场,还会削弱其被上下文引导至未受训练偏好视角的能力。实验将模型微调至文化价值观争议的某一方后,受训立场在日常使用中愈发主导,而识别并忠实执行对立视角的能力持续下降。作者提出按指定视角分布最大化奖励的替代目标,并给出立场分布匹配(stance-distribution matching)作为实现方案。

正文

View PDF HTML (experimental)

Abstract:AI models serving a heterogeneous population must act on the principles appropriate to each user and context. While post-training has been shown to narrow the views large language models express, prior work has focused on default behavior rather than the ability to adapt to in-context information. We show that post-training also degrades a model's ability to be steered in-context toward perspectives it was not trained to favor. In controlled experiments, we fine-tune models toward one side of cultural-value disagreements and evaluate checkpoints throughout training. The trained side becomes increasingly dominant in ordinary use, while the ability to recognize and faithfully enact the opposing view declines. These findings point to a tension between prioritizing a single set of values and preserving the technical capacity needed to serve diverse stakeholders. Finally, we propose and analyze an alternative objective that maximizes reward subject to a prescribed distribution over expressed perspectives, and present stance-distribution matching as a practical implementation.
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2610.02614 [cs.LG]
  (or arXiv:2610.02614v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.02614

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Itai Shapira [view email]
[v1] Fri, 2 Oct 2026 00:14:40 UTC (616 KB)

来源:arXiv:cs.LG · arxiv.org