arXiv:cs.LG· Narek Maloyan·· 3 小时前AI 评分44
评估并提升大语言模型对输入序列变化的鲁棒性
Evaluating and Improving the Robustness of Large Language Models to Input Sequence Variations
AI 导读
一篇博士论文提出评估与提升 LLM 对抗输入序列变化鲁棒性的模型、方法与算法,包括基于 Jensen-Shannon 散度的生成鲁棒性指标 R_stab(f)。
正文
Abstract:Large language models (LLMs) in production systems face prompt injections, trojans (backdoors), and manipulation of automatic quality metrics. This thesis develops models, methods, and algorithms for evaluating and improving LLM robustness to adversarial input sequence variations. We propose R_stab(f), a generative robustness metric based on the Jensen-Shannon divergence between per-step output distributions under small input perturbations. For localized attacks we prove V(h) <= 1 - R_class(h), where R_class(h) is the probability that a decision operator h keeps its decision under small perturbations. For non-localized attacks we propose a calibrated empirical model. For LLM-as-a-Judge systems we develop ASA, an adaptive evolutionary black-box attack that reaches an attack success rate (ASR) of up to 73.8%, with transfer between open models up to 62.6%. On Trojan Detection Challenge 2023 data (Pythia-1.4B), surrogate triggers reach REASR ~0.99 while recall of the true triggers is ~0.17 against a baseline of ~0.14. On SaTML CTF 2024 we systematize four classes of bypasses of multi-layer defenses, which reduce the ASR from 90% to 15-25%. Committees of 5-7 heterogeneous models reduce the ASR for Gemma-3-4B by 47-55 percentage points, to 19.3% with 7 models. For agentic systems based on the Model Context Protocol (MCP), we propose AttestMCP, which attests tool calls with HMAC-protected packets at under 0.1 ms per call, and the Commit Boundary isolation pattern. On the MCPBench benchmark of 847 scenarios they reduce the average ASR from 53.7% to 12.4%. The methods are implemented in the JudgeGuard and TrojanArmor software suites and the MCPSec module.
| Comments: | PhD thesis, 2026. 118 pages |
| Subjects: | Cryptography and Security (cs.CR); Computation and Language (cs.CL); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.02432 [cs.CR] |
| (or arXiv:2610.02432v1 [cs.CR] for this version) | |
| https://doi.org/10.48550/arXiv.2610.02432 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Narek Maloyan [view email]
[v1]
Thu, 1 Oct 2026 19:57:56 UTC (154 KB)
来源:arXiv:cs.LG · arxiv.org