arXiv:cs.LG· Eliyahu Levy, Adam Teman, Yoni Pugachov·· 4 小时前AI 评分38
TR-PTQ:通过 Taylor Region 重构实现高精度纯整数 Transformer 训练后量化
TR-PTQ: High-Accuracy Integer-Only Transformer Post Training Quantization via Taylor Region Reformulation
AI 导读
TR-PTQ 是一种纯整数 Transformer 训练后量化方法,通过共享 Taylor Region 指数与对数原语,让除法和平方根等运算完全在 log 域内用标准整数算术完成。研究指出归一化层的学习缩放参数与 GELU 的复合近似是主要误差来源,而 SoftMax 对激进量化本身鲁棒。该方法无需浮点硬件单元处理非线性层,在视觉与语言基准上绝对精度损失低于 1.5%。
正文
Abstract:Post-training quantization (PTQ) enables efficient deployment, yet transformer architectures remain challenging to quantize due to nonlinear layers. While existing methods attribute accuracy loss to insufficient numerical precision, often necessitating floating-point fallbacks, we demonstrate that degradation is actually driven by specific structural error sources. We find that learned scale parameters in normalization layers and compounded approximations in GELU are the primary error contributors, whereas SoftMax remains inherently robust to aggressive quantization. To address these bottlenecks, we introduce TR-PTQ, a unified integer-only formulation using shared Taylor Region (TR) exponential and logarithm primitives. This approach allows computationally expensive operations, including division and square roots, to be performed entirely in the log-domain via standard integer arithmetic. Combined with a calibration-free, outlier-aware optimization for LayerNorm parameters, our method eliminates the need for floating-point hardware units for nonlinearities, achieving less than 1.5\% absolute accuracy degradation across vision and language benchmarks.
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.09969 [cs.LG] |
| (or arXiv:2610.09969v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.09969 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Eliyahu Levy [view email]
[v1]
Wed, 7 Oct 2026 12:39:10 UTC (1,111 KB)
来源:arXiv:cs.LG · arxiv.org