How Differently Do Memory Faults Affect Quantized Lightweight Neural Networks?
Abstract
The explosive growth in the size of neural networks is driving an equally explosive growth in the memory needed to support inferencing, such as weights. This growth brings an increase in the possibility of memory errors sneaking through error-correction checkpoints and perturbing weights. In addition, it has also motivated contributors to build lightweight ML frameworks to deploy such models on edge devices efficiently. Yet, this does not count as a leeway to consider the system to be more fault-tolerant. Non-determinism in critical environments with lightweight neural networks, such as embedded networks in a space radiation-filled environment, remains an obvious issue and raises questions of reliability; overlooking a largely unaddressed question is thus: How do such errors and bit-flips in weights impact these pruned neural networks' outputs? As volatility and unreliable accuracy can be an obvious answer, our work aims to explore a comprehensive relationship between the characteristics that shape the fault and the architecture that defines the neural network. Hence, this paper reports on a first systemic study of what happens to a pruned TensorFlow Lite (TFLite) neural network under a combination of different relevant cases, and the often near-catastrophic effect they may have on network results.