jietang· @jietang · X·· 2026-08-21AI 评分46
AI 导读
精彩评论:FLOPs 是智能;参数是知识!
正文
cool comment: FLOPs were intelligence; parameters were knowledge!
An excellent history of scaling laws from @jietang. In 2020, we explored the limits of sparsity in Switch Transformers by routing each token to only 1 out of 2048 experts (in retrospect, a bold choice). The model had fewer than 3B activated parameters, but 1.6T total parameters (comparable to today's frontier models). The 1.6T model achieved better C4 perplexities than the T5 models using far less compute, set a new SOTA on TriviaQA, but was dumb as bricks on reasoning tasks like SuperGLUE. The lesson was that the optimal tokens-per-parameter ratio is highly task-dependent. Or as @NShazeer had already intuited: FLOPs were intelligence; parameters were knowledge!在 X 查看被引用的帖子
来源:jietang · x.com