跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Yohan Chatelain (Krembil Centre for Neuroinformatics, CAMH, Toronto, Canada), Pablo de Oliveira Castro (Universite Paris-Saclay, UVSQ, LI-PaRAD, Versailles, France)·· 17 小时前AI 评分41

低精度 Transformer 推理该用随机舍入还是就近舍入?基于小型 GPT-2 的可变精度仿真研究

Stochastic Rounding in Low-Precision Transformer Inference: A Variable-Precision Emulation Study of a Small GPT-2

AI 导读

研究将 PRISM 向量化舍入库扩展为可变精度随机舍入(VPSR)算法,在固定数值格式下只改变各运算位置的舍入规则,对 DistilGPT-2 进行低精度推理仿真。

正文

Authors:Yohan Chatelain (1), Pablo de Oliveira Castro (2) ((1) Krembil Centre for Neuroinformatics, CAMH, Toronto, Canada, (2) Universite Paris-Saclay, UVSQ, LI-PaRAD, Versailles, France)

View PDF

Abstract:Should low-precision transformer inference use stochastic rounding (SR) or round-to-nearest (RN)? The answer depends on where in the network you look. We isolate this effect by holding the numerical format fixed and varying only the rounding rule at individual operation sites. To enable experiments at freely chosen precisions, we extend the PRISM vectorized rounding library to arbitrary virtual precision via a variable-precision stochastic rounding (VPSR) algorithm, proving that the rounding decision is evaluated exactly in hardware floating point.
We develop two analyses providing complementary insight into this site-level trade-off. First, a probabilistic forward-error bound for linear projections shows that SR's error envelope grows as $O(\sqrt{n} u)$ in reduction length $n$, versus $O(n u)$ for RN, a gap that widens rapidly at low precision and is most pronounced in the long multilayer perceptron (MLP) down-projection. Second, a second-order decomposition of expected cross-entropy loss change at the output softmax into signed drift, drift curvature, and a Fisher-weighted variance penalty reveals why the two sites behave oppositely: MLP noise is predominantly a uniform logit shift to which softmax is invariant, so SR's variance is largely discounted; head noise is non-uniform across the vocabulary and is not.
On DistilGPT-2 at $t=6$ significand bits, observations match theory: SR in the MLP raises perplexity to 1.15x the full-precision reference, versus 2.21x for RN. At the language-model head, the ordering reverses because SR introduces non-uniform variance, whereas deterministic RN carries none. In a mixed-precision configuration (MLP output at $t=6$), assigning SR to the MLP and RN to the head brings perplexity within 1.10x of the full-precision reference, a 28% reduction over matched-bit RN.
Comments: 35 pages, 10 figures, 4 tables. Code and evaluation pipeline available at this https URL and archived on Zenodo at this https URL
Subjects: Machine Learning (cs.LG); Computation and Language (cs.CL)
ACM classes: G.1.0; I.2.6; I.2.7
Cite as: arXiv:2610.01889 [cs.LG]
  (or arXiv:2610.01889v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.01889

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Pablo De Oliveira Castro [view email]
[v1] Thu, 1 Oct 2026 15:40:48 UTC (413 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org