arXiv:cs.LG· Xueqi Cheng, Liang Wu, Kelly Wan, Liangjie Hong, Yushun Dong·· 3 小时前AI 评分41
LLM 压缩的能力缩放下降定律
Capability Scaling-Down Laws for LLM Compression
AI 导读
研究系统考察了剪枝、量化与蒸馏三类 LLM 压缩方法的能力缩放下降规律,在数学、代码生成和问答任务上测量能力损失,并建立可预测的简单关系式。跨剪枝层级共享密度响应可将拟合剪枝预测器所需的配置测量量减半,在 Pythia 新状态、预注册的 OLMo-2 测试状态及 Wanda 剪枝下,该紧凑关系与全量测量回归在数学和代码上的差距在 0.020 nats per token 以内。
正文
Abstract:LLM compression reduces inference costs and memory requirements, but selecting a method and configuration remains largely empirical because comparable resource reductions can produce different capability losses. We systematically investigate capability scaling-down laws for LLM compression across pruning, quantization, and distillation. Our framework measures capability loss in mathematics, code generation, and question answering, and relates these measurements to model size, training stage, compression settings, data availability, and training exposure. We develop simple predictive relations and evaluate their accuracy, measurement efficiency, and generalization to unseen configurations and model states. Sharing the density response across pruning levels halves the configuration measurements needed to fit a pruning predictor: on new Pythia states, on pre-registered OLMo-2 test states and under Wanda pruning, the compact relation matches a regression fitted with all measurements on math and code to within 0.020 nats per token, with coefficients refitted for each setting. Controlled distillation experiments show that the cost of heavy data reuse recurs across question-answering distributions, while the net benefit depends on the evaluation distribution. We further evaluate the decision value of these predictions by comparing numerical selection with configuration medians and fixed method priorities. Independent evaluations across two model families show that selection captures most of the available cross-method benefit for question answering within the tested candidate sets, where a fixed method priority attains the same regret, with smaller opportunities for mathematics and code. These results clarify the predictive scope of capability scaling-down laws and their use in compression method selection. Our code is publicly available at: this https URL.
| Subjects: | Machine Learning (cs.LG); Computation and Language (cs.CL) |
| Cite as: | arXiv:2610.02462 [cs.LG] |
| (or arXiv:2610.02462v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.02462 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Xueqi Cheng [view email]
[v1]
Thu, 1 Oct 2026 20:35:12 UTC (1,477 KB)
来源:arXiv:cs.LG · arxiv.org