跳到正文
arXiv:cs.LG· Mart\'in Bravo, Samuel Horv\'ath, Gonzalo Navarro, Andr\'es Abeliuk·· 3 小时前AI 评分36

NeuralZip:可复用的快速无损压缩方案

NeuralZip: Reusable Setup for Fast Lossless Compression

AI 导读

NeuralZip 通过将指数分布相近的块分组、共享 Huffman 码并选择性打包重复指数元组,实现可复用的无损压缩设置。在浮点模型检查点上,设置后压缩速度比基线快 1.81-21.33×,并实现精确的逐比特重建,该设置还可从兼容架构预计算并迁移。GPU 实验中活跃内存占用最多降低 27.5%,且 logits 完全一致。

正文

View PDF

Abstract:Lossless compression can reduce the storage and movement of model weights without changing their floating-point values, but repeated statistical analysis and code construction add computational overhead. We study whether the statistical structure of exponents can be prepared once and reused. For this, we introduce NeuralZip, which groups chunks with similar exponent distributions, shares Huffman codes, and selectively represents recurring exponent tuples using packed exponents, thereby achieving additional moderate compression ratios. A setup chooses these representations before subsequent encodings, while every encoding still processes the current tensor values. In floating-point model checkpoints, post-setup compression is 1.81-21.33$\times$ faster than the baselines and achieves exact bit-to-bit reconstruction. We show that this setup can be precomputed and transferred from another compatible architecture, preserving similar compression ratios and avoiding the need to amortize setup costs. Therefore, compression adaptation is transferable and reusable. Training checkpoints demonstrate continued reuse as the weights evolve. Finally, GPU experiments reduce active memory usage by up to 27.5$\%$ while reproducing the logits exactly.
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2610.09916 [cs.LG]
  (or arXiv:2610.09916v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.09916

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Martín Bravo [view email]
[v1] Wed, 7 Oct 2026 12:06:24 UTC (883 KB)

来源:arXiv:cs.LG · arxiv.org