跳到正文
arXiv:cs.AI· Tianyi Wang, Huawei Fan, Yuanchao Shu, Peng Cheng, Cong Wang·· 3 小时前

重新审视延迟拒绝服务攻击:攻击 LLM 服务框架而非模型

Rethinking Latency Denial-of-Service: Attacking the LLM Serving Framework, Not the Model

AI 导读

研究发现算法层面的延迟攻击对现代 LLM 服务系统基本无效,连续批处理等系统级优化可隔离延迟影响。作者转而提出 Fill and Squeeze 攻击策略,针对调度器状态转换,先在 vLLM 上耗尽全局 KV cache 引发队头阻塞,再迫使系统反复抢占。在 vLLM 上实测 TTFT 最高退化 75-742 倍,TPOT 平均减速 1.5-4 倍,攻击成本比现有方法低 30-40%。

正文

View PDF HTML (experimental)

Abstract:LLM inference is inherently expensive, even a modest slowdown can translate into substantial operating costs and severe availability risks. Recently, a growing body of research known as latency attacks focuses on crafting inputs to trigger worst-case output lengths. However, we report a contrary finding that these algorithmic-level latency attacks are largely ineffective against modern LLM serving systems. We reveal that system-level optimization such as continuous batching provides a logical isolation to mitigate contagious latency impact on co-located users. Thus, in this paper, we shift our focus from the algorithm to the system layer, and introduce a new Fill and Squeeze attack strategy targeting the state transition of the scheduler. ``Fill'' first exhausts the global KV cache to induce Head-of-Line blocking, while ``Squeeze'' forces the system into repetitive preemption. By manipulating output lengths using different attack prompts, and leveraging side-channel probing of memory status, we demonstrate that the attack can succeed in a practical black-box setting with much less cost. Extensive evaluations on vLLM indicate up to $75-742\times$ TTFT degradation relative to benign baselines and $1.5-4\times$ average slowdown on Time Per Output Token compared to existing attacks with 30-40% lower attack cost. Code: this https URL
Comments: NeurIPS 2026
Subjects: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
Cite as: arXiv:2602.07878 [cs.CR]
  (or arXiv:2602.07878v2 [cs.CR] for this version)
  https://doi.org/10.48550/arXiv.2602.07878

arXiv-issued DOI via DataCite

Submission history

From: Tianyi Wang [view email]
[v1] Sun, 8 Feb 2026 09:05:54 UTC (9,554 KB)
[v2] Thu, 8 Oct 2026 07:33:40 UTC (10,308 KB)

来源:arXiv:cs.AI · arxiv.org