跳到正文
arXiv:cs.AI· Yuyang Dai, Rana Shahout, Mahmood Sharif·· 5 小时前AI 评分47

JIL:利用长度预测漏洞攻击 LLM 调度器

Jumping the Line: Exploiting Length Predictions in LLM Scheduling

AI 导读

研究者提出 JIL,一种针对基于预测的 LLM 调度器的攻击,通过优化对抗后缀使轻量级输出长度探针低估请求长度,从而获取更高调度优先级。在两种数据集和四个 LLM 上,JIL 将预测输出长度最多降低 83.4%,对抗请求端到端平均完成速度提升最多 1.53 倍。将长度预测分组为粗粒度区间可削弱 JIL 的调度优势并缓解对正常请求的延迟。

正文

View PDF HTML (experimental)

Abstract:Efficient request scheduling is increasingly important for reducing completion time in large language model (LLM) serving. Size-based policies such as Shortest Job First prioritize shorter requests, but output lengths are unknown before generation, so practical schedulers rely on predicted lengths. We introduce JIL, an attack on prediction-based LLM schedulers that manipulates the scheduling signal to obtain higher priority and reduce completion time. Using TRAIL as a case study, JIL optimizes an adversarial suffix that causes a lightweight output-length probe to underestimate a request's length. We evaluate JIL on two datasets and four LLMs across varied request profiles and deployment configurations. JIL reduces predicted output lengths by up to 83.4 percent, and adversarial requests complete up to 1.53 times faster on average in end-to-end serving experiments. The reduction in predicted length is substantially larger than the change in actual output length, revealing a mismatch between the scheduler's estimate and the request's realized size. Response utility varies across models and tasks, exposing a trade-off between scheduling advantage and response quality. We also evaluate scheduler-side defenses and find that grouping length predictions into coarse intervals reduces JIL's scheduling advantage and mitigates delays to benign requests.
Comments: 24 pages, 6 figures
Subjects: Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.03430 [cs.AI]
  (or arXiv:2610.03430v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.03430

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Yuyang Dai [view email]
[v1] Fri, 2 Oct 2026 15:12:43 UTC (218 KB)

来源:arXiv:cs.AI · arxiv.org