跳到正文
arXiv:cs.LG· Shiwen Wang, Pengxiang Zhao, Xiaoming Yuan·· 3 小时前

Deflating Hessian:面向多模态扩散 Transformer 的秩-4 W4A4 量化

Deflating the Hessian: Rank-4 W4A4 Quantization for Multimodal Diffusion Transformers

AI 导读

研究者提出 Deflating Hessian(\method{})框架,将低秩辅助的 W4A4 训练后量化建模为耦合校准问题,通过"收缩 Hessian"扣除低秩分量已捕获的残差误差,并引入激活噪声代理抑制激活量化误差。

正文

View PDF HTML (experimental)

Abstract:In diffusion transformers, low-rank branches can mitigate 4-bit weight--activation (W4A4) post-training quantization (PTQ) loss by decomposing each weight into a low-bit residual and a high-precision low-rank component. Existing low-rank PTQ approaches, however, either optimize low-rank compensation and residual quantization separately, often requiring higher ranks, or rely on second-order weight updates without explicitly modeling activation quantization error, which becomes particularly pronounced under 4-bit quantization. To address these limitations, we present \method{}, a unified framework modeling low-rank-assisted W4A4 PTQ as a coupled calibration problem and deriving optimization-based solvers from the joint objective. Eliminating the output-side low-rank factor yields a \emph{deflated Hessian} that discounts residual errors already captured by the low-rank component, while an activation-noise surrogate is incorporated to suppress activation quantization error. Across five diffusion backbones, rank-4 \method{} consistently outperforms rank-4 SVDQuant in PSNR and LPIPS. It further surpasses rank-32 SVDQuant on SANA-1.6B, FLUX.1-schnell, and FLUX.1-dev with an $8\times$ smaller rank and up to $6.25\times$ faster quantization. Furthermore, on the Qwen3-8B LLM, rank-4 \method{} improves MMLU accuracy from 61.50\% to 68.17\% over rank-32 SVDQuant. Overall, \method{} achieves better W4A4 performance with substantially lower rank and quantization cost.
Subjects: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV); Optimization and Control (math.OC)
Cite as: arXiv:2610.11315 [cs.LG]
  (or arXiv:2610.11315v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.11315

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Shiwen Wang [view email]
[v1] Thu, 8 Oct 2026 06:22:19 UTC (3,416 KB)

来源:arXiv:cs.LG · arxiv.org