跳到正文
原文
Together AI 研究与产品博客(RSS)·· 2026-05-11精选AI 评分68

Together AI 解析 DeepSeek-V4 服务化:百万 token 上下文为何是推理系统工程问题

Serving DeepSeek-V4: why million-token context is an inference systems problem

AI 导读

Together AI 发文解析 DeepSeek-V4 的服务化挑战,认为其核心变化是在 token 轴压缩 KV cache 的混合注意力设计(CSA、HCA、SWA),使模型支持 1M token 上下文窗口。

推荐理由

原文基于 Together 在 HGX B200 上的实际 bring-up,给出 KV cache 布局和缓存策略如何决定 DeepSeek-V4 长上下文吞吐的具体经验。

来源:Together AI 研究与产品博客(RSS) · together.ai