Right In-Place Convolution: A Simple, General, and Near-Optimal Strategy for Memory-Efficient CNN Inference
Opegbemi M. Busoye ⋅ Tolulope Busoye ⋅ Eghonghon-aye Eigbe
Abstract
Activation memory, not compute, limits CNN inference on the microcontrollers that carry most edge deployments in low-resource settings. Direct in-place convolution removes the dual-buffer cost, but the memory-optimal formulation of Gural and Murmann [2019] assumes valid padding, unit stride, unit dilation and odd square kernels, and needs a non-sequential traversal costing $2\times$ inference time. We identify two regimes in which their published closed form does not hold: an under-allocation of exactly $(k-1)C_{in} \bmod (C_{out}-C_{in})$ scalars, active on every convolutional layer of their own deployed network, and an unbounded overestimate, up to $2{,}432\times$, once the critical leg leaves the output grid. We correct both and generalize to arbitrary stride, dilation, padding and rectangular kernels. We then propose Right In-Place (RiP) convolution, a bit-identical operation in which every layer reads its input right-aligned in a shared workspace and writes its output left-aligned from index zero, with a debt piecewise affine in the output pixel index so that nine breakpoints give the minimum safe gap in $O(1)$. Across 84 layers from 25 architectures RiP matches the herringbone workspace exactly on 58 and within 5\% on 81. Written into TinyEngine's kernels and deployed to a Raspberry Pi Pico 1 and Pico 2, it cuts peak activation memory across eleven MCUNet models by 12.5 to 33.3\% at unchanged cycle counts and bit-identical outputs, taking the Pico 1 from six of the eleven models to nine.
Chat is not available.
Successful Page Load