Understanding the Convergence of Direct Training of SNNs with Surrogate Gradients
Abstract
Spiking Neural Networks (SNNs) offer high energy efficiency, but direct optimization is challenging because spike generation is governed by a non-differentiable hard threshold. Practical direct-training methods, including surrogate gradients (SG) and local zeroth-order (LocalZO) estimators, maintain hard spikes in the forward pass while using surrogate or perturbation-induced derivatives in the backward pass. This forward--backward mismatch raises a central theoretical question: how do local gradient discrepancies propagate through the spatiotemporal computation graph, and can the resulting biased updates still provide guaranties for the original discrete-spike objective? We develop a unified convergence analysis framework for SG and LocalZO direct SNN training. A smoothed spike objective is introduced only as an analytical bridge, allowing us to separate the core error into two components: the gradient-level mismatch between the direct-training mean direction and the smoothed-objective gradient, and the objective-level consistency gap between the smoothed and original discrete objectives. Both components are controlled by a threshold-tail functional that measures how often hard-trajectory membrane potentials lie near the firing threshold. Under suitable regularity assumptions, we obtain a non-asymptotic guarantee for the original discrete-spike objective. Experiments across representative SNN architectures and benchmarks verify convergence consistency, the existence of errors, and demonstrate that the threshold-tail statistic is an effective diagnostic for direct SNN training.