跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Robert Hu·· 14 小时前AI 评分41

面向快速 FP4 预训练的格式感知融合方法

Format-Aware Fusion for Fast FP4 Pretraining

AI 导读

研究提出"格式感知融合"方法,将每个量化生产者与其 scale 域和消费者布局协同设计,覆盖原生 mxfp、全局 nvfp 和 CTA-local nvfp。

正文

View PDF HTML (experimental)

Abstract:Four-bit floating-point (FP4) Tensor Cores accelerate matrix multiplication, but scale computation, operand packing, layout construction, and saved backward state can erase the gain. We present \emph{format-aware fusion}, which co-designs each quantization producer with its scale domain and consumer layout for native \mxfp{}, global \nvfp{}, and cooperative-thread-array-local \nvfp{}. We evaluate Llama-3-family 8B pretraining through 160 billion tokens using bfloat16 output projections and compiled cross entropy. In matched same-accelerator probes, bfloat16 and Transformer Engine \nvfp{} reach 18.8K and 27.6K tokens/s/GPU, while our fastest custom route reaches 37.9K. \mxfp{} with row-gradient stochastic rounding and fixed-sign 32-value Hadamard weight-gradient preconditioning reaches 37.2K tokens/s/GPU (86.3\% bfloat16 model FLOP utilization) and ends 2.11\% above the raw bfloat16 training-loss endpoint. A Transformer Engine recipe with four final bfloat16 blocks ends 0.87\% above bfloat16 at 27.1K tokens/s/GPU. Downstream rankings differ from training-loss rankings, showing that FP4 outcomes depend jointly on scale contract, operand, and execution path.
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2610.00053 [cs.LG]
  (or arXiv:2610.00053v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.00053

arXiv-issued DOI via DataCite

Submission history

From: Robert Hu [view email]
[v1] Fri, 4 Sep 2026 00:42:51 UTC (2,481 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org