跳到正文
arXiv:cs.AI· Kartik Ramesh, Kaidi Fu, Zihan Zheng, Jiahuan Yu, Fabio Oliveira, Carlos Costa, Minjia Zhang·· 4 小时前AI 评分44

FluidPD:面向 SLO 的预填充-解码分离式 LLM 服务原地弹性调度

FluidPD: In-Place Elasticity for SLO-Aware Prefill-Decode Disaggregated LLM Serving

AI 导读

FluidPD 是一个面向 SLO 的 P/D 分离式 LLM 服务系统,通过 FluidToken 在解码侧空闲时卸载部分预填充计算、FluidRole 原地切换运行中 worker 的预填充与解码角色,应对短时突发与持续性的阶段负载失衡。

正文

View PDF HTML (experimental)

Abstract:Prefill-decode disaggregation is becoming a common architecture for LLM serving because it separates two phases with distinct execution patterns and SLO objectives. Existing systems typically combine a fixed prefill/decode worker ratio with request routing across workers. However, real-world workloads exhibit both short bursts and sustained shifts in the prefill-to-decode demand ratio. As a result, a configuration that is well provisioned at one time may quickly become mismatched, causing latency SLO violations even when idle capacity exists elsewhere. Existing autoscaling mechanisms can add capacity, but they react slowly, require spare GPUs, and do not directly address short-timescale phase imbalance.
We present FluidPD, a P/D-disaggregated serving system that provides SLO-aware in-place elasticity. FluidPD introduces two complementary mechanisms. FluidToken handles transient imbalance by offloading a bounded portion of prefill computation to decode workers when decode-side slack is available. FluidRole handles sustained imbalance by reassigning running workers between prefill and decode roles in place, avoiding model reload and engine restart. Both mechanisms are guided by lightweight pressure indices that expose prefill and decode-side resource pressure before they appear as SLO violations. Across production Azure trace workloads, FluidPD improves overall SLO attainment over static SGLang by up to 94.6 percentage points, demonstrating that SLO-aware in-place P/D elasticity improves service quality without provisioning additional workers.
Comments: 13 pages, 11 figures
Subjects: Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.06917 [cs.AI]
  (or arXiv:2610.06917v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.06917

arXiv-issued DOI via DataCite

Submission history

From: Kaidi Fu [view email]
[v1] Fri, 2 Oct 2026 18:49:43 UTC (1,043 KB)

来源:arXiv:cs.AI · arxiv.org