跳到正文
arXiv:cs.LG· Rui Liu, Benjamin Paa{\ss}en·· 3 小时前AI 评分37

W16A16:在微控制器上以 8-bit 成本实现 16-bit 精度卷积神经网络

16-bit Precision of Convolutional Neural Networks on Microcontroller Units for 8-bit Costs

AI 导读

研究提出 W16A16 量化方法,在 Armv7E-M 微控制器上以 16-bit 精度实现比 8-bit 量化低约 10 倍的量化误差,同时保持相近或更优的推理时间与能耗。该方法在层级和模型级均比替代量化方案更快、更省电,并针对回归与分类任务评估了实际量化误差及 MCU 部署中的时间与能耗表现。

正文

View PDF HTML (experimental)

Abstract:To deploy deep neural networks on edge hardware, highly efficient inference schemes are necessary that retain high accuracy. This work presents W16A16, a high precision (16-bit), fast speed, low energy quantization method. On a widely applied microcontroller architecture Armv7E-M, our proposed approach achieves faster speed and lower energy consumption on layer- and model-level compared to alternative quantization schemes. We analyze the architecture of Armv7E-M, explain the underlying principles behind the performance advantages of 16-bit approaches, and evaluate the empiric quantization errors for regression and classification tasks, as well as empiric time- and energy consumption in MCU deployment. We observe ca.\ 10 times lower quantization errors compared to 8-bit quantization schemes while achieving similar or better inference times and energy consumption.
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2610.03402 [cs.LG]
  (or arXiv:2610.03402v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.03402

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Benjamin Paassen [view email]
[v1] Fri, 2 Oct 2026 14:50:26 UTC (564 KB)

来源:arXiv:cs.LG · arxiv.org