LMSYS:Blog(Chatbot Arena 团队)·· 16 小时前AI 评分45
SGLang、Qwen 与 NVIDIA 团队如何用 NVFP4 KV Cache 加速长上下文与智能体推理
Blog Accelerating Long-Context and Agentic Inference with NVFP4 KV Cache The KV cache is a fundamental building block of the modern LLM inference system. The context from multiple conversation rounds in agent sessions is cached as keys and values (KV) in GPU memory, allowi... SGLang, Qwen, and NVIDIA teams September 16, 2026
AI 导读
SGLang、Qwen 和 NVIDIA 团队在 SGLang 中实现了 NVFP4 KV cache,将 K/V 以 4 位 NVFP4 格式存储并在 decode 阶段于注意力 kernel 内即时反量化到 FP8。
来源:LMSYS:Blog(Chatbot Arena 团队) · lmsys.org