跳到正文
arXiv:cs.LG· Abbas Raza Ali, Muhammad Ajmal Siddiqui, Moona Zahid·· 10 小时前AI 评分35

面向企业工作负载的生成式 AI 推理算力评估框架

Evaluating Inference Compute for Generative AI: A Framework for Enterprise Workloads

AI 导读

一项 arXiv 研究提出面向企业工作负载的生成式 AI 推理算力评估框架,指出 LLM 部署正从单轮补全转向智能体轨迹,使推理硬件经济性反转。

正文

View PDF HTML (experimental)

Abstract:LLM deployment is shifting from single-turn completion to agentic trajectories in which a model plans, calls tools, reads results and reasons at test time before acting. This inverts the economics of inference hardware: chat serving amortises weight reads across large batches, whereas agent trajectories are sequentially dependent, run at effective batch one, and make per-token decode latency (TPOT) the dominant term in task completion time. Using a roofline analysis and a closed-form episode-latency model, we show why this regime favours accelerators that keep weights in on-die SRAM (Cerebras WSE-3/3T, Groq/NVIDIA LPU) or compiler-managed tiered memory (SambaNova SN40L/SN50), and why three vendor ecosystems converged in 2026 on disaggregated prefill/decode serving. We show that per-step reliability compounds exponentially in trajectory length-a 2% per-step failure rate erases a 2x decode advantage for a 20-step agent-so determinism and tail latency are first-order performance variables. We then propose a four-layer evaluation framework (silicon, serving system, agent episode, enterprise) with a metric set built on goodput at an agentic SLO and cost per successful episode, a six-axis benchmark protocol over six task families, a paired-bootstrap statistical design, an attestation protocol for vendor-run benchmarks, and TCO, availability and adoption-timing models with explicit break-even conditions. All performance figures are public and labelled by evidence class; we state seven falsifiable hypotheses and the experiments that test them, and argue that the most likely original result is that token-throughput rankings diverge from cost-per-successful-task rankings on long-horizon work.
Subjects: Hardware Architecture (cs.AR); Distributed, Parallel, and Cluster Computing (cs.DC); Machine Learning (cs.LG); Performance (cs.PF)
Cite as: arXiv:2610.07094 [cs.AR]
  (or arXiv:2610.07094v1 [cs.AR] for this version)
  https://doi.org/10.48550/arXiv.2610.07094

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Abbas Raza Ali [view email]
[v1] Mon, 5 Oct 2026 13:31:06 UTC (254 KB)

来源:arXiv:cs.LG · arxiv.org