跳到正文
arXiv:cs.LG· JuneHyung Kim, Sankeerth Durvasula, Nandita Vijaykumar·· 3 小时前AI 评分40

SpAx:用权重近似与激活稀疏加速卸载权重的 LLM 解码

Activation Sparsity with Weight Approximation for Faster LLM Decoding on Offloaded Weights

AI 导读

SpAx 将激活稀疏中"是否读取权重"的二元选择改为三选项:完全保留、用压缩权重近似、或直接跳过,从而在权重卸载到 CPU 内存时平均加速解码 3.86 倍(最高 5.57 倍,16-bit 权重),WikiText-2 困惑度增幅不超过 10%。权重卸载到闪存时平均加速 3.31 倍(最高 4.81 倍),4-bit 权重下分别为 2.06 倍和 1.54 倍。

正文

View PDF HTML (experimental)

Abstract:Deploying LLMs on consumer-grade GPUs with insufficient memory to hold their weights can result in prohibitively slow inference, because decoding repeatedly transfers offloaded weights from system RAM or flash storage into GPU at much lower bandwidth than local GPU-memory access. Activation sparsity reduces these transfers by skipping weights associated with zero or near-zero activations. However, as more activation contributions are omitted, model quality eventually degrades rapidly, indicating that weights associated with small-magnitude activations collectively influence model quality sharply. In this work, we improve the trade-off between model quality and decoding performance when exploiting activation sparsity. Our key idea is to replace the binary choice of whether or not to read a weight with three options: fully retain it, approximate it using a compressed weight representation, or omit it entirely. SpAx skips weights associated with activations closest to zero, reads approximate weights for smaller-magnitude activations, and reads original weights for the largest-magnitude activations. Smaller-magnitude activations attenuate the errors introduced by approximate weights, while compressed weight representations require fewer bytes to be transferred. With weights offloaded to CPU memory, SpAx speeds up decoding by 3.86X on average (up to 5.57X) with 16-bit weights and 2.06X (up to 2.74X) with 4-bit weights, at a WikiText-2 perplexity increase of at most 10%. With weights offloaded to flash storage, the speedups are 3.31X on average (up to 4.81X) and 1.54X (up to 2.03X).
Subjects: Machine Learning (cs.LG); Performance (cs.PF)
Cite as: arXiv:2610.02598 [cs.LG]
  (or arXiv:2610.02598v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.02598

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: JuneHyung Kim [view email]
[v1] Thu, 1 Oct 2026 23:50:52 UTC (408 KB)

来源:arXiv:cs.LG · arxiv.org