arXiv:cs.LG(机器学习,全量分类)· Robert Hu·· 14 小时前AI 评分41
面向快速 FP4 预训练的格式感知融合方法
Format-Aware Fusion for Fast FP4 Pretraining
AI 导读
研究提出"格式感知融合"方法,将每个量化生产者与其 scale 域和消费者布局协同设计,覆盖原生 mxfp、全局 nvfp 和 CTA-local nvfp。
正文
Abstract:Four-bit floating-point (FP4) Tensor Cores accelerate matrix multiplication, but scale computation, operand packing, layout construction, and saved backward state can erase the gain. We present \emph{format-aware fusion}, which co-designs each quantization producer with its scale domain and consumer layout for native \mxfp{}, global \nvfp{}, and cooperative-thread-array-local \nvfp{}. We evaluate Llama-3-family 8B pretraining through 160 billion tokens using bfloat16 output projections and compiled cross entropy. In matched same-accelerator probes, bfloat16 and Transformer Engine \nvfp{} reach 18.8K and 27.6K tokens/s/GPU, while our fastest custom route reaches 37.9K. \mxfp{} with row-gradient stochastic rounding and fixed-sign 32-value Hadamard weight-gradient preconditioning reaches 37.2K tokens/s/GPU (86.3\% bfloat16 model FLOP utilization) and ends 2.11\% above the raw bfloat16 training-loss endpoint. A Transformer Engine recipe with four final bfloat16 blocks ends 0.87\% above bfloat16 at 27.1K tokens/s/GPU. Downstream rankings differ from training-loss rankings, showing that FP4 outcomes depend jointly on scale contract, operand, and execution path.
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.00053 [cs.LG] |
| (or arXiv:2610.00053v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.00053 arXiv-issued DOI via DataCite |
Submission history
From: Robert Hu [view email]
[v1]
Fri, 4 Sep 2026 00:42:51 UTC (2,481 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org