跳到正文
arXiv:cs.CL· Wenjun Wang, Heng Li, Yanggan Gu, Hongxia Yang·· 3 小时前AI 评分35

OnlineQAT:面向超低比特大语言模型的在线策略蒸馏

OnlineQAT: On-Policy Distillation for Ultra-Low-Bit Large Language Models

AI 导读

OnlineQAT 是一个两阶段框架,先通过分块 QAT 获得可用的低比特初始化,再对学生模型自生成回答做在线策略蒸馏(OPD),由冻结的全精度教师模型在每个前缀处提供采样的 reverse-KL 训练信号。在 Qwen3-1.7B 上,W3A16 平均分 57.28、W2A16 为 32.52,分别比 ReasoningQAT 提升 2.90 和 0.44 分,三比特下收益尤为明显。

正文

View PDF HTML (experimental)

Abstract:Quantization-aware training (QAT) can recover much of the accuracy lost when large language models are compressed below four bits. Existing re- covery stages, however, are commonly optimized on fixed completions or teacher-generated answers, whereas the deployed quantized model condi- tions on prefixes generated by itself. Quantization errors can therefore move the model into states that are absent from offline recovery data. We introduce OnlineQAT, a two-stage framework that first obtains a usable low-bit initialization through block-wise QAT and then performs on-policy distillation (OPD) on student-generated responses. At each visited pre- fix, a frozen full-precision teacher provides a sampled reverse-KL training signal. On Qwen3-1.7B, OnlineQAT obtains the best average among the compared quantized methods: 57.28 at W3A16 and 32.52 at W2A16, im- proving over ReasoningQAT by 2.90 and 0.44 points, respectively. The results suggest that student-visited states provide a useful recovery signal beyond fixed-completion training, particularly at three bits.
Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
Cite as: arXiv:2610.09346 [cs.CL]
  (or arXiv:2610.09346v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2610.09346

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Wenjun Wang [view email]
[v1] Wed, 7 Oct 2026 03:05:42 UTC (66 KB)

来源:arXiv:cs.CL · arxiv.org