arXiv:cs.LG· Joe Dwyer·· 4 小时前AI 评分25
缩减缩放定律:资源受限大语言模型中的参数效率与计算最优训练
Scaling Down the Scaling Laws: Parameter Efficiency and Compute-Optimal Training in Resource-Constrained Large Language Models
AI 导读
一篇综述梳理了 LLM 缩放理论从经验缩放定律到计算最优训练的演进,重点关注参数效率、token 利用率、数据效率与资源受限环境,涵盖数据剪枝、高效架构、量化、低秩适配与面向边缘的优化。文章指出研究正从规模最大化转向对参数、token、算力与硬件的审慎分配,并认为未来进展应以性能、参数量、计算成本、token 分配与硬件约束之间的关系来评估效率,而非仅看模型表现。
正文
Abstract:Large language models (LLMs) have achieved substantial performance gains through increases in model size, training data, and computational resources. However, traditional scaling approaches produce diminishing returns, rising financial and environmental costs, and barriers to participation for researchers operating outside large industrial laboratories. This review examines the evolution of LLM scaling theory from empirical scaling laws to compute-optimal training, with particular emphasis on parameter efficiency, token utilization, data efficiency, and resource-constrained environments. Foundational work on scaling laws is synthesized alongside later research on compute-optimal training, data pruning, efficient architectures, quantization, low-rank adaptation, and edge-oriented optimization. The literature indicates a shift from scale maximization toward more deliberate allocation of parameters, tokens, compute, and hardware resources. At the same time, important empirical, theoretical, and methodological gaps remain regarding whether scaling principles established on enterprise-grade infrastructure generalize to smaller models and constrained computing environments. This review organizes these developments into a unified framework for resource-efficient LLM training and argues that future progress should evaluate efficiency not solely through model performance, but through the relationship among performance, parameter count, computational cost, token allocation, and hardware constraints.
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.06387 [cs.LG] |
| (or arXiv:2610.06387v2 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.06387 arXiv-issued DOI via DataCite |
Submission history
From: Joseph Dwyer [view email]
[v1]
Mon, 5 Oct 2026 14:12:53 UTC (23 KB)
[v2]
Wed, 7 Oct 2026 17:24:21 UTC (23 KB)
来源:arXiv:cs.LG · arxiv.org