跳到正文
原文
Rohan Paul· @rohanpaul_ai · X·· 4 小时前AI 评分32
AI 导读

Cerebras 联合创始人兼 CEO Andrew Feldman 解释,其晶圆级架构在 LLM 推理中比 GPU 快 2500 倍。推理分 pre-fill 和 decode 两阶段,decode 阶段每生成一个 token 都需将模型权重从内存搬入计算单元;GPU 从 HBM 搬运,而 Cerebras 将权重存放在分布在整个晶圆级处理器上的 SRAM 中,使这一搬运速度快约 2500 倍。

正文

Andrew Feldman, co-founder and CEO of Cerebras gives the best explanation of why Cerebras' wafer-scale architecture is 2,500X faster than a GPU during LLM inference.

During inference, there are 2 stages:
- pre-fill, where the model first processes the user's prompt, and
- decode, where it generates the answer 1 token at a time in sequence.

During that sequencial Decode phase, before each token is calculated the model weights have to be moved from memory into compute.

On a GPU those weights are moved from HBM, while Cerebras keeps them in much faster SRAM spread across its very large wafer-scale processor, so the memory-to-compute movement that must happen for every token is about 2,500× faster

----
From The MAD Podcast with Matt Turck and Cerebras YouTube channel, (link in comment)

来源:Rohan Paul · x.com