跳到正文
arXiv:cs.AI· Fan Huang, Haewoon Kwak, Jisun An·· 5 小时前AI 评分44

LLM 道德推理轨迹研究:基于探针的可解释性探索

Understanding Moral Reasoning Trajectories in Large Language Models: Toward Probing-Based Explainability

AI 导读

研究者提出"道德推理轨迹"概念,即 LLM 中间推理步骤中伦理框架调用的序列,并在 6 个模型和 3 个基准上分析其动态。55.4–57.7% 的连续步骤发生框架切换,仅 16.4–17.8% 的轨迹保持框架一致;不稳定轨迹遭受说服攻击的概率高 1.29 倍。

正文

View PDF HTML (experimental)

Abstract:Large language models (LLMs) increasingly participate in morally sensitive decision-making, yet how they organize ethical frameworks across reasoning steps remains underexplored. We introduce moral reasoning trajectories, sequences of ethical framework invocations across intermediate reasoning steps, and analyze their dynamics across six models and three benchmarks. We find that moral reasoning involves systematic multi-framework deliberation: 55.4--57.7% of consecutive steps involve framework switches, and only 16.4--17.8% of trajectories remain framework-consistent. Unstable trajectories remain 1.29 times more susceptible to persuasive attacks (p=0.015). At the representation level, linear probes localize framework-specific encoding to model-specific layers (layer 63/81 for Llama-3.3-70B; layer 17/81 for Qwen2.5-72B), achieving 16.8--22.2% lower KL divergence than the step-prior baseline. Activation steering applied during generation moves the framework-consistency--accuracy relationship, widening it for Qwen2.5-72B and erasing it for Llama-3.3-70B, and a probe-space layer sweep bounds the attainable drift reduction at 6.7--8.9%. We further propose a Moral Representation Consistency (MRC) metric whose underlying framework attributions are validated by human annotators (mean cosine similarity = 0.859), and we report what an automated coherence rater does and does not establish about it.
Comments: We updated some more statistical results and analysis
Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
Cite as: arXiv:2603.16017 [cs.CL]
  (or arXiv:2603.16017v2 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2603.16017

arXiv-issued DOI via DataCite

Submission history

From: Fan Huang [view email]
[v1] Mon, 16 Mar 2026 23:51:30 UTC (7,844 KB)
[v2] Mon, 5 Oct 2026 20:29:09 UTC (6,984 KB)

来源:arXiv:cs.AI · arxiv.org