跳到正文
原文
Together AI 研究与产品博客(RSS)·· 2026-03-04AI 评分41

Together AI 提出 CPD 缓存感知预填充-解码分离架构,长上下文 LLM 服务吞吐提升最高 40%

Cache-aware prefill–decode disaggregation (CPD) for up to 40% faster long-context LLM serving

AI 导读

Together AI 提出缓存感知的预填充-解码分离架构 CPD,在延迟 SLO 下可持续 QPS 比现有分离式设计提升 35–40%,并在存在大型冷提示时保持更紧的尾延迟。

来源:Together AI 研究与产品博客(RSS) · together.ai