arXiv:cs.LG(机器学习,全量分类)· Anudeex Shetty, Aditya Joshi, Salil S. Kanhere·· 15 小时前AI 评分46
用“醉酒语言”诱导 LLM 安全失效:人格提示、因果微调与强化后训练
In Vino Veritas and Vulnerabilities: Examining LLM Safety via Drunk Language Inducement
AI 导读
研究提出通过人格提示、因果微调和强化后训练三种机制,在 LLM 中诱导“醉酒语言”,即模拟酒精影响下写出的文本。在 5 个 LLM 上评测发现,这类模型在 JailbreakBench 上更易被越狱(即使存在防御),在 ConfAIde 上也更易泄露隐私,且效果优于基线模型和此前已报告的方法。
正文
Abstract:Humans are susceptible to undesirable behaviours and privacy leaks under the influence of alcohol. This paper investigates drunk language, i.e., text written under the influence of alcohol, as a driver for safety failures in large language models (LLMs). We investigate three mechanisms for inducing drunk language in LLMs: persona-based prompting, causal fine-tuning, and reinforcement-based post-training. When evaluated on 5 LLMs, we observe a higher susceptibility to jailbreaking on JailbreakBench (even in the presence of defences) and privacy leaks on ConfAIde, where both benchmarks are in English, as compared to the base LLMs as well as previously reported approaches. Via a robust combination of manual evaluation and LLM-based evaluators and analysis of error categories, our findings highlight a correspondence between human-intoxicated behaviour, and anthropomorphism in LLMs induced with drunk language. The simplicity and efficiency of our drunk language inducement approaches position them as potential counters for LLM safety tuning, highlighting significant risks to LLM safety.
| Comments: | Accepted to INLG 2026 |
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR); Machine Learning (cs.LG) |
| Cite as: | arXiv:2601.22169 [cs.CL] |
| (or arXiv:2601.22169v2 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2601.22169 arXiv-issued DOI via DataCite |
Submission history
From: Anudeex Shetty [view email]
[v1]
Mon, 19 Jan 2026 12:44:20 UTC (7,873 KB)
[v2]
Thu, 1 Oct 2026 12:14:18 UTC (7,837 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org