arXiv:cs.LG· Arnold Olympio, Juan Manuel Servera Bondroit, Wael Abdelmalek, Guang Lu, Jo\~ao Carvalho·· 4 小时前AI 评分36
可复现 LLM 推理基准测试:面向回归测试的顺序隔离协议
Reproducible LLM Inference Benchmarking: A Sequential Isolation Protocol for Regression Testing
AI 导读
研究提出 Sequential Isolation Methodology,一种通过受控基准测试与回归测试协议降低 LLM 推理测量方差的方法。
正文
Abstract:Reproducible benchmarking of Large Language Model (LLM) inference is challenging because repeated measurements can vary with execution and system state. We present the Sequential Isolation Methodology, a controlled benchmarking and regression-testing protocol designed to reduce between-run measurement variance while deliberately varying workload concurrency. We evaluate three representative open-source LLMs on an NVIDIA A100 80GB GPU using vLLM 0.9.1 across six context sizes and eight concurrency levels, with five repetitions per configuration. The final protocol reduces average coefficient of variation (CV) from 15.2% in the least controlled methodology stage to 2.2% under the final protocol; using CV computed across the five repetition-level median (P50) TTFT values per configuration, 113 of 144 configurations (78.5%) achieve CV below 3%. The measurements also show a marked latency transition between 200 and 500 concurrent users on the tested stack and descriptive differences in P99 latency across the three models. We additionally provide an explicit cost break-even model with sensitivity to API pricing. The protocol is intended to provide a stable reference for reproducible comparison and regression testing rather than to predict absolute behavior under uncontrolled production traffic. Infrastructure-as-Code and benchmark scripts support replication of the experimental environment.
| Subjects: | Performance (cs.PF); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.09778 [cs.PF] |
| (or arXiv:2610.09778v1 [cs.PF] for this version) | |
| https://doi.org/10.48550/arXiv.2610.09778 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Guang Lu [view email]
[v1]
Wed, 7 Oct 2026 09:58:30 UTC (3,552 KB)
来源:arXiv:cs.LG · arxiv.org