Lilian Weng· @lilianweng · X·· 2026-06-26AI 评分46
AI 导读
一篇拖了很久(3 年多?)的缩放定律文章。 算力很贵。缩放定律帮我们在投入大规模训练之前,推理数据与模型规模之间的最优算力分配。 文章涵盖:缩放定律预测了什么、算力最优分配如何运作、Kaplan et al. 与 Chinchilla 为何结论不一致,以及数据限制 + 拟合细节如何让外推变得棘手。 https://lilianweng.github.io/posts/2026-06-24-scaling-laws/
正文
A super long overdue (3+ years?) post on scaling laws.
Compute is expensive. Scaling laws are a way to help us reason about the optimal compute allocation between data and model size before committing to a large run.
The post covers what scaling laws predict, how compute-optimal allocation works, why Kaplan et al. and Chinchilla disagree, and how data limits + fitting details make extrapolation tricky.
https://lilianweng.github.io/posts/2026-06-24-scaling-laws/
来源:Lilian Weng · x.com