跳到正文
arXiv:cs.LG· Orion Reblitz-Richardson·· 4 小时前AI 评分61

arXiv 论文:后训练决定 LLM 是否会违背自身道德判断行事

Principled Under Pressure: Post-Training Decides Whether LLMs Act on Their Own Moral Judgment

AI 导读

一篇 arXiv 论文(arXiv:2610.08670)构建了覆盖五种压力的 248 个预注册场景,让同一模型分别以 Agent 身份选择行动和以第三人称判断对错,以模型自身判断为参照。

正文

View PDF HTML (experimental)

Abstract:Language models increasingly act as agents. An agent that says an action is wrong and then takes it anyway is a different failure from one that does not know better, and evaluations of stated values cannot see it. We build a pre-registered panel of 248 scenarios across five kinds of pressure. Each scenario is posed twice to the same model, once as the agent choosing what to do and once in the third person asking which option is right, so the model's own judgment is the reference. Every scenario has a twin with the pressure removed, and every model gets a positive control in which its operator orders the violating action, so that a missing gap can be told apart from a blind instrument. On OLMo-3-7B-Instruct, the model takes the action it judged wrong on about one in five pressuring scenarios, more often than on the same scenarios with the pressure removed. Across four instruct models the gap depends on the post-training recipe: OLMo-3 and Meta's Llama-3.1-8B-Instruct carry it; Tulu 3 shows none on the whole panel (above about 0.01 in probability) or on its own most-pressuring scenarios; Qwen2.5-7B-Instruct shows none on the whole panel (above about 0.02) and is unresolved on its own (0.083, -0.028 to 0.195). Meta's recipe and Ai2's Tulu 3 start from the same Llama-3.1 weights, and only Meta's carries the gap. Reading a chat model outside its chat template reverses the sign of its gap with nothing at stake (-0.038 against +0.055 under the template on OLMo-3), a distortion present on two of three recipes. On both models that carry it, reasoning about the stakes before acting moves the choice back toward the model's own judgment, against a same-length non-moral task, with or without the pressure; on OLMo-3, naming the norm at stake does about a third of that. The gap is a measurable target for post-training recipes, not a fixed property of pretrained weights.
Comments: 33 pages
Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
Cite as: arXiv:2610.08670 [cs.LG]
  (or arXiv:2610.08670v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.08670

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Orion Reblitz-Richardson [view email]
[v1] Tue, 6 Oct 2026 16:52:23 UTC (281 KB)

来源:arXiv:cs.LG · arxiv.org