arXiv:cs.LG(机器学习,全量分类)· Opegbemi Matthias Busoye, Tolulope Matthew Busoye, Eghonghon-aye Eigbe·· 14 小时前AI 评分41
Right In-Place (RiP) 卷积:一种简单、通用且近最优的内存高效 CNN 推理策略
Right In-Place (RiP) Convolution: A Simple, General, and Near-Optimal Strategy for Memory-Efficient CNN Inference
AI 导读
研究者提出 Right In-Place (RiP) 卷积,通过让每层在共享工作区右对齐读取输入、从索引 0 左对齐写出输出,实现位级一致的 CNN 推理。
正文
Abstract:Activation memory, not compute, limits CNN inference on constrained hardware such as microcontrollers. Direct in-place convolution removes the dual-buffer cost, but the memory-optimal formulation of Gural and Murmann assumes valid padding, unit stride, unit dilation, and odd square kernels, and needs a non-sequential traversal costing $2\times$ inference time in transposes. We identify two regimes in which their published closed form does not hold: (1) an under-allocation of exactly $(k-1)C_{in} \bmod (C_{out}-C_{in})$ scalars, active on every convolutional layer of their own deployed network and manifesting as a silent corruption of still-live input; (2) an unbounded overestimate, up to $2{,}432\times$, once the critical leg leaves the output grid. We correct both and generalize to arbitrary stride, dilation, padding, and rectangular kernels. We then propose Right In-Place (RiP) convolution, a bit-identical operation in which every layer reads its input right-aligned in a shared workspace and writes its output left-aligned from index zero. The debt is piecewise affine in the output pixel index, so evaluating its breakpoints in $O(1)$ yields the minimum safe gap without enumerating the output grid, with row-major access preserved. Across $10{,}000$ random layers RiP produced no corruption, and across 84 convolutional layers from 25 architectures it matches the herringbone workspace exactly on 58 and within 5% on 81, using 24.8% less memory than dual buffering on average. Written into TinyEngine's kernels and deployed to a Raspberry Pi Pico 1 and Pico 2, it cuts peak activation memory across eleven MCUNet models by 12.5 to 33.3% at unchanged cycle counts and bit-identical outputs, raising the number of models that fit the Pico 1's 256 KB SRAM from six to nine.
| Comments: | Extended version of a paper accepted at the NeurIPS 2026 Workshop on Global South in AI |
| Subjects: | Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV) |
| ACM classes: | C.3; I.4.0; D.4.2; B.3.2 |
| Cite as: | arXiv:2610.00586 [cs.LG] |
| (or arXiv:2610.00586v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.00586 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Opegbemi Busoye Mr. [view email]
[v1]
Wed, 30 Sep 2026 18:51:30 UTC (164 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org