跳到正文
原文
Fuli Luo· @_LuoFuli · X·· 2026-05-30AI 评分47
AI 导读

小米 MiMo-V2.5 与 MiMo-V2.5-Pro 基于 Hybrid SWA 架构,将 KVCache 存储压缩至 Full Attention 的约 1/7。

正文

Inference Optimizations Behind the MiMo-V2.5 Series API Price Reductions

Read the full technical blog: https://mimo.xiaomi.com/blog/mimo-v2-5-inference

The V2.5 model family, including MiMo-V2.5 and MiMo-V2.5-Pro, is built on a Hybrid Sliding Window Attention (Hybrid SWA) architecture, which compresses KVCache storage to roughly 1/7 that of Full Attention. However, architectural advantages rarely translate directly into measurable gains in production serving. To realize these gains, we redesigned KVCache management, tiered caching, and the prefix-cache tree; addressed key challenges in SWA KVCache handling; and optimized scheduling as well as the Prefill/Decode pipeline.

Validated on real production traffic, these optimizations have increased effective KVCache capacity by nearly 5x, with server-side cache hit rates averaging 93%–95% across mainstream harness frameworks. Together with MoE configuration tuning and multimodal inference optimizations, they enable more efficient long-context inference and form part of what makes the recent API price cuts possible.

来源:Fuli Luo · x.com