arXiv:cs.LG· Abbas Raza Ali, Muhammad Ajmal Siddiqui, Moona Zahid·· 10 小时前AI 评分35
面向企业工作负载的生成式 AI 推理算力评估框架
Evaluating Inference Compute for Generative AI: A Framework for Enterprise Workloads
AI 导读
一项 arXiv 研究提出面向企业工作负载的生成式 AI 推理算力评估框架,指出 LLM 部署正从单轮补全转向智能体轨迹,使推理硬件经济性反转。
正文
Abstract:LLM deployment is shifting from single-turn completion to agentic trajectories in which a model plans, calls tools, reads results and reasons at test time before acting. This inverts the economics of inference hardware: chat serving amortises weight reads across large batches, whereas agent trajectories are sequentially dependent, run at effective batch one, and make per-token decode latency (TPOT) the dominant term in task completion time. Using a roofline analysis and a closed-form episode-latency model, we show why this regime favours accelerators that keep weights in on-die SRAM (Cerebras WSE-3/3T, Groq/NVIDIA LPU) or compiler-managed tiered memory (SambaNova SN40L/SN50), and why three vendor ecosystems converged in 2026 on disaggregated prefill/decode serving. We show that per-step reliability compounds exponentially in trajectory length-a 2% per-step failure rate erases a 2x decode advantage for a 20-step agent-so determinism and tail latency are first-order performance variables. We then propose a four-layer evaluation framework (silicon, serving system, agent episode, enterprise) with a metric set built on goodput at an agentic SLO and cost per successful episode, a six-axis benchmark protocol over six task families, a paired-bootstrap statistical design, an attestation protocol for vendor-run benchmarks, and TCO, availability and adoption-timing models with explicit break-even conditions. All performance figures are public and labelled by evidence class; we state seven falsifiable hypotheses and the experiments that test them, and argue that the most likely original result is that token-throughput rankings diverge from cost-per-successful-task rankings on long-horizon work.
| Subjects: | Hardware Architecture (cs.AR); Distributed, Parallel, and Cluster Computing (cs.DC); Machine Learning (cs.LG); Performance (cs.PF) |
| Cite as: | arXiv:2610.07094 [cs.AR] |
| (or arXiv:2610.07094v1 [cs.AR] for this version) | |
| https://doi.org/10.48550/arXiv.2610.07094 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Abbas Raza Ali [view email]
[v1]
Mon, 5 Oct 2026 13:31:06 UTC (254 KB)
来源:arXiv:cs.LG · arxiv.org