arXiv:cs.CL· Or Biton, Tomer Krichli, Itai Allouche, Joseph Keshet·· 4 小时前AI 评分45
请求中隐藏的线索:用 Token 相关性解释 LLM 的不道德顺从
Hidden in the Request: Explaining Unethical LLM Compliance through Token Relevance
AI 导读
研究用客观分类、第一人称陈述和直接求助三种形式向 LLM 呈现不道德场景,发现模型在"求助式"表述下表现最差。借助 Layer-wise Relevance Propagation(LRP),作者将这一差异归因于归因偏差:模型更关注"能帮我个忙吗"这类良性任务框架 token,而忽视"别被抓到"这类暗示不道德行为的 cue-token。
正文
Abstract:Although Large Language Models (LLMs) are aligned to optimize for both helpfulness and harmlessness, these dual objectives may conflict, inevitably leading to alignment failures. This work systematically investigates instances where LLMs fail to exhibit ethical behavior. To understand the underlying mechanics of these vulnerabilities, we introduce a probing methodology that presents unethical scenarios to LLMs in three distinct structural modalities: objective classification tasks, subjective first-person statements, and direct requests for assistance. We find that model performance degrades in the request-for-assistance-based form. Using Layer-wise Relevance Propagation (LRP), we trace this discrepancy to an attribution bias: the model places greater emphasis on benign task-framing tokens (e.g., "Can you help me...") than on tokens signaling the underlying unethical behavior (e.g., "without getting caught"), which we term cue-tokens. We hypothesize that this under-attribution contributes to harmful compliance. To test this, we introduce two LRP-guided decoding methods that steer generation toward trajectories more relevant to cue tokens. Empirical evaluations show that these interventions promote safer responses, supporting cue-token attribution's role in compliance failures.
| Comments: | SocialAgent, NeurIPS 2026 |
| Subjects: | Artificial Intelligence (cs.AI); Computation and Language (cs.CL) |
| Cite as: | arXiv:2608.23264 [cs.AI] |
| (or arXiv:2608.23264v3 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2608.23264 arXiv-issued DOI via DataCite |
|
| Journal reference: | NeurIPS 2026, SocialAgent Workshop |
Submission history
From: Itai Allouche [view email]
[v1]
Mon, 24 Aug 2026 13:52:05 UTC (45 KB)
[v2]
Mon, 5 Oct 2026 10:27:22 UTC (45 KB)
[v3]
Tue, 6 Oct 2026 18:36:13 UTC (44 KB)
来源:arXiv:cs.CL · arxiv.org