跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Opegbemi Matthias Busoye, Tolulope Matthew Busoye, Eghonghon-aye Eigbe·· 14 小时前AI 评分41

Right In-Place (RiP) 卷积:一种简单、通用且近最优的内存高效 CNN 推理策略

Right In-Place (RiP) Convolution: A Simple, General, and Near-Optimal Strategy for Memory-Efficient CNN Inference

AI 导读

研究者提出 Right In-Place (RiP) 卷积,通过让每层在共享工作区右对齐读取输入、从索引 0 左对齐写出输出,实现位级一致的 CNN 推理。

正文

View PDF HTML (experimental)

Abstract:Activation memory, not compute, limits CNN inference on constrained hardware such as microcontrollers. Direct in-place convolution removes the dual-buffer cost, but the memory-optimal formulation of Gural and Murmann assumes valid padding, unit stride, unit dilation, and odd square kernels, and needs a non-sequential traversal costing $2\times$ inference time in transposes. We identify two regimes in which their published closed form does not hold: (1) an under-allocation of exactly $(k-1)C_{in} \bmod (C_{out}-C_{in})$ scalars, active on every convolutional layer of their own deployed network and manifesting as a silent corruption of still-live input; (2) an unbounded overestimate, up to $2{,}432\times$, once the critical leg leaves the output grid. We correct both and generalize to arbitrary stride, dilation, padding, and rectangular kernels. We then propose Right In-Place (RiP) convolution, a bit-identical operation in which every layer reads its input right-aligned in a shared workspace and writes its output left-aligned from index zero. The debt is piecewise affine in the output pixel index, so evaluating its breakpoints in $O(1)$ yields the minimum safe gap without enumerating the output grid, with row-major access preserved. Across $10{,}000$ random layers RiP produced no corruption, and across 84 convolutional layers from 25 architectures it matches the herringbone workspace exactly on 58 and within 5% on 81, using 24.8% less memory than dual buffering on average. Written into TinyEngine's kernels and deployed to a Raspberry Pi Pico 1 and Pico 2, it cuts peak activation memory across eleven MCUNet models by 12.5 to 33.3% at unchanged cycle counts and bit-identical outputs, raising the number of models that fit the Pico 1's 256 KB SRAM from six to nine.
Comments: Extended version of a paper accepted at the NeurIPS 2026 Workshop on Global South in AI
Subjects: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
ACM classes: C.3; I.4.0; D.4.2; B.3.2
Cite as: arXiv:2610.00586 [cs.LG]
  (or arXiv:2610.00586v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.00586

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Opegbemi Busoye Mr. [view email]
[v1] Wed, 30 Sep 2026 18:51:30 UTC (164 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org