跳到正文
arXiv:cs.AI· Neriah Ben David, Ori Meir, Or Ordentlich·· 6 小时前AI 评分36

OSFP4:面向 NVFP4 量化的对角平滑与块缩放联合优化

OSFP4: Joint Optimization of Diagonal Smoothing and Block Scales for NVFP4 Quantization

AI 导读

研究者提出 OSFP4,一种针对 NVFP4 量化的新方案,为每个线性投影引入对角平滑矩阵,并将其与块缩放进行联合优化以最小化 NVFP4 下的矩阵乘积量化误差。

正文

View PDF HTML (experimental)

Abstract:NVFP4 is an attractive datatype for large language model (LLM) inference, offering compact storage and native tensor-core acceleration. However, preserving accuracy using NVFP4 requires careful quantization. In this work we develop a novel quantization scheme called Optimized Smoothing and Scaling for NVFP4 (OSFP4). For each linear projection it uses a diagonal smoothing matrix whose entries are optimized to minimize the squared matrix-product quantization error under NVFP4, taking into account the rounding procedure that is used (either round-to-nearest, or GPTQ-style successive interference cancellation). This requires performing joint optimization on the smoothing entries as well as the block scales, which is facilitated by analyzing a multiplicative-dither FP4 quantizer instead of the fixed deterministic one. Experiments show that OSFP4 achieves the highest average accuracy among the evaluated competitors in the corresponding quantization settings, while retaining approximately 94-97\% of vendor NVFP4 prefill throughput on the measured workloads. Our code is available in this https URL
Subjects: Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.08231 [cs.AI]
  (or arXiv:2610.08231v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.08231

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Or Ordentlich [view email]
[v1] Tue, 6 Oct 2026 12:17:52 UTC (656 KB)

来源:arXiv:cs.AI · arxiv.org