跳到正文
arXiv:cs.AI· Zhixu Du, Weijia Han, Hai Helen Li, Yiran Chen·· 5 小时前AI 评分44

一个 Token 需要多少算力?用 Mixture-of-Agents 测量充分逐 Token 计算量

What Does a Token Cost? A Mixture-of-Agents Measurement of Sufficient Per-Token Compute

AI 导读

研究用 Mixture-of-Agents(MoA)方法测量单个 token 实际所需的计算量:由来自三个模型族的十五个不同容量模型逐 token 复现参考序列,以最小成功智能体的推理成本作为该 token 的充分算力。

正文

View PDF HTML (experimental)

Abstract:Large language models spend the same amount of computation on every token they generate, regardless of how difficult each token is to produce. Methods such as speculative decoding and model routing are built on the premise that much of this computation is unnecessary, yet the computation an individual token actually requires has not been measured. We measure it through a Mixture-of-Agents (MoA) lens: a panel of fifteen language models of increasing capacity, drawn from three families, in which every agent attempts to reproduce a reference sequence token by token, conditioned on the correct preceding tokens. We define the inference cost of the smallest agent that succeeds as the token's sufficient compute, which upper-bounds what the token requires. On three core benchmarks, a 0.5B agent reproduces 92--95\% of reference tokens. Across Qwen, OLMo, and R1-distilled panels, the most expensive 10\% account for 64--80\% of estimated FLOPs. On all 500 MATH-500 problems, the MoA-derived map helps model routing reduce projected latency from 7.59 to 5.12 seconds while slightly improving accuracy, relative to the best confidence-routing baseline. The MoA-map helps drafting use 32.6\% fewer draft tokens and approximately 20\% lower projected latency than fixed-window drafting at similar accuracy. These comparisons reveal remaining allocation headroom, motivating controllers that exploit sufficient-compute structure.
Subjects: Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.02491 [cs.AI]
  (or arXiv:2610.02491v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.02491

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Zhixu Du [view email]
[v1] Thu, 1 Oct 2026 21:11:41 UTC (163 KB)

来源:arXiv:cs.AI · arxiv.org