arXiv:cs.LG· Haoyu Zheng, Fangcheng Fu, Binhang Yuan, Yongqiang Zhang, Liang Deng, Hao Wang, Yuanyuan Zhu, Xiao Yan, Jiawei Jiang·· 2 天前AI 评分39
Denoising Surface:为 Diffusion LLM 服务建模与预测推理成本
Denoising Surface: Modeling and Predicting Inference Cost for Diffusion LLM Serving
AI 导读
研究者提出 Denoising Workload Surface(DWS),将 diffusion LLM 的块自回归生成结构保留为二维“块-去噪步”概率曲面,用于加权各步异构成本。
正文
Abstract:As diffusion large language models (dLLMs) become more capable, they are moving from research settings to real-world \textit{serving}, where request management (such as scheduling and resource allocation) relies on accurate estimation of per-request inference cost. However, common cost proxies fall short for dLLMs: output length ignores that one forward pass can unmask multiple tokens, and denoising-step count ignores the \textit{heterogeneous} per-step costs. We observe that the block-autoregressive generation mechanism induces a two-dimensional execution structure over output blocks and within-block denoising steps, whereas these proxies collapse it into a scalar, discarding information essential for characterizing the cost. Motivated by this insight, we propose the Denoising Workload Surface (DWS), which preserves this two-dimensional block-step structure as a probability surface to weight the heterogeneous per-step costs. We then design a coarse-to-fine training scheme that enables a lightweight prompt-only predictor to accurately predict the complex DWS. This predictor runs efficiently even on a single CPU core, avoiding GPU contention with the serving model. Since DWS decouples request-dependent execution behavior from deployment-specific cost factors, the predictor transfers across hardware configurations without retraining. In \textit{real-world} serving experiments, DWS reduces cost-prediction error by up to $2.50\times$ over scalar-based predictors, while the DWS-guided shortest-job-first scheduler reduces end-to-end latency by up to $1.92\times$ for online chatbots.
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.00499 [cs.LG] |
| (or arXiv:2610.00499v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.00499 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Haoyu Zheng [view email]
[v1]
Wed, 30 Sep 2026 18:01:37 UTC (885 KB)
来源:arXiv:cs.LG · arxiv.org