arXiv:cs.LG· Neeraj Karamchandani, Piyush Nagasubramaniam, Xinhong Xie, Sencun Zhu, Dinghao Wu·· 5 小时前AI 评分50
arXiv 论文提出 TPRS:Agent 安全基准分数随威胁表述变化显著波动
Threat-Preserving Representation Sensitivity in Agent-Security Benchmarks
AI 导读
论文提出威胁保持表示敏感性(TPRS),衡量在任务、策略与评测标准不变时改变 Agent 可见表示对攻击成功率(ASR)的影响。
正文
Abstract:Security benchmarks for LLM-based agents often report the attack success rate (ASR) as a measure of model robustness and use these scores to compare different models and defense mechanisms, assuming that they describe the security of the agent. In this paper, we explore whether it also influences the benchmark's measurement.
To measure the effect of the benchmark representation, we introduce threat-preserving representation sensitivity (TPRS), which measures how much the ASR changes when we change the agent-visible representation while holding the underlying task, harmful action, security policy, ground truth, environment, and the evaluation criteria fixed.
On Agent Security Bench (ASB), replacing threat-related tool names with threat-neutral names raises the committed attack success rate by 11.67 percentage points on GPT-5-mini and by 13.21 points on Claude Haiku 4.5. On MCPTox, replacing the original neutral tool name with an explicit threat-related name lowers the ASR by 11.00 percentage points on GPT-5-mini and 4.11 points on Claude Haiku 4.5. On AgentDojo, adding threat-related wording to the attack-relevant tool changes ASR by only 0.50 percentage points on GPT-4o-mini, yet the benign utility falls by 5.36 points on tasks requiring that tool.
We ran an experiment on MCPTox where we observed that a threat-neutral name matched on token count, length, and casing reproduces most of the shift produced by the threat-explicit name (8.54 of 11.00 points on GPT-5-mini).
The results show that a security score measured under one representation may fail to generalize across threat-preserving representations of the same security problem. Robustness claims should therefore be supported by performance across a controlled set of threat-preserving representations rather than relying on a single representation-dependent score.
| Comments: | 12 pages, 2 figures |
| Subjects: | Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.03585 [cs.CR] |
| (or arXiv:2610.03585v1 [cs.CR] for this version) | |
| https://doi.org/10.48550/arXiv.2610.03585 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Neeraj Karamchandani [view email]
[v1]
Fri, 2 Oct 2026 16:55:50 UTC (257 KB)
来源:arXiv:cs.LG · arxiv.org