跳到正文
arXiv:cs.LG· Donghyun Lee, Arkapravo Ghosh, Varun Manjunath, Bumjoon Kyle Rhee, Hyunho Kook, Shiting Xiao, Youngeun Kim, Priyadarshini Panda·· 4 小时前AI 评分37

SoloQ:面向扩散语言模型的无校准量化框架

SoloQ: Calibration-Free Quantization for Diffusion Language Models

AI 导读

针对扩散语言模型(dLLM)因全序列去噪和高推理成本难以部署的问题,研究者提出无校准量化框架 SoloQ,通过将权重和激活映射到具有可预测边缘分布的归一化旋转基,实现与数据无关的量化,并结合结构化 K-RPBH 旋转与轻量重缩放校正。

正文

View PDF HTML (experimental)

Abstract:Diffusion large language models dLLMs) have emerged as a promising alternative to autoregressive language models through bidirectional diffusion-based token generation. However, their growing model sizes and high inference costs make efficient deployment challenging: full-sequence denoising repeatedly invokes compute-intensive forward passes, while block-diffusion models additionally introduce a memory-intensive KV-cache. Low-bit weight-activation quantization is therefore attractive, yet existing dLLM post-training quantization methods rely on calibration data despite activation distributions shifting across masking states and denoising steps. We present SoloQ, a calibration-free quantization framework that maps weights and activations into a normalized rotated basis with a predictable marginal distribution, enabling data-independent quantization. SoloQ combines a structured K-RPBH rotation with a lightweight rescaling correction for calibration-free quantization. Its predictable post-rotation distribution supports both distribution-matched codebooks and hardware-native NVFP4. For block-diffusion models, SoloQ further applies commit-time KV-cache quantization to compress persistent states without perturbing the actively denoised block. Across full-sequence dLLMs (LLaDA and Dream) and block-diffusion dLLMs(Fast-dLLM v2 and Nemotron-Labs-Diffusion), SoloQ retains accuracy under 4-bit quantization and outperforms calibration-based baselines on knowledge- and reasoning-intensive benchmarks. With NVFP4, SoloQ reduces peak memory by up to 2.61X and accelerates end-to-end inference by up to 2.24X.
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2610.07121 [cs.LG]
  (or arXiv:2610.07121v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.07121

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Donghyun Lee [view email]
[v1] Mon, 5 Oct 2026 17:30:54 UTC (2,854 KB)

来源:arXiv:cs.LG · arxiv.org