跳到正文
arXiv:cs.CL· Haitong Jiang, Chunlin Liu, Sihan Tang, Chan Wu, Xiaoqing Su, Yuhong Feng·· 4 小时前

LLM 安全评估中的意图恢复:相同结果,不同证据

Same Outcome, Different Evidence: Intent Recovery in LLM Safety Evaluation

AI 导读

针对 LLM 安全评估普遍只用攻击成功率(ASR)汇总有害输出行为的问题,研究者提出将 ASR 与操作理解率(UR)配对报告,以区分模型是识别并拒绝了有害任务、未能识别任务,还是答非所问。受控英文重构实验显示,提示词越明确,意图恢复率越高,而 ASR 并不遵循同样规律。代码与实验输入已公开。

正文

View PDF HTML (experimental)

Abstract:Safety evaluations of large language models commonly summarize harmful-output behavior with attack success rate (ASR). Yet the same non-harmful outcome can arise for very different reasons. A model may recover a harmful task and refuse it, fail to recover the task, or respond to something else entirely. Distinguishing these cases becomes especially important under intent-obscuring prompts, where a low ASR does not reveal whether the evaluated task was actually engaged. To make this distinction explicit, we pair ASR with operative understanding rate (UR), which measures whether a response both identifies the evaluated task and treats it as the task to be answered. Across interfaces, this paired view reveals substantial variation hidden by ASR: similar ASR values can correspond to sharply different recovery rates. Controlled English reconstructions show that recovery consistently improves as compressed prompts become more explicit, whereas ASR does not follow the same pattern. A complementary contrast comes from FormalLogic, where high recovery can still coincide with frequent harmful assistance. Together, these results show that non-harmful outcomes are not equally informative about model safety, motivating the joint reporting of intent recovery and ASR in LLM safety evaluation. Code and experiment inputs are available at this https URL.
Comments: 12 pages, 2 figures. Code and experiment inputs: this https URL
Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.11766 [cs.CL]
  (or arXiv:2610.11766v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2610.11766

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Haitong Jiang [view email]
[v1] Thu, 8 Oct 2026 11:50:32 UTC (302 KB)

来源:arXiv:cs.CL · arxiv.org