Memory-Regularized Forward–Reverse Layer-wise Optimization
Abstract
We ask whether a deep network can be optimized through independently solvable layer-local prob- lems while remaining aligned with a global supervised objective. Recursive Layer-wise Regression (RLR) and Recursive Layer-wise Classification (RLC) answer this question with a forward–reverse procedure: a forward sweep records layer inputs, a reverse sweep converts supervision into desired preactivations, and each affine layer solves a regularized least-squares problem. For classification, normalized classifier directions and distribution-agnostic noise produce diverse hidden targets. To reduce minibatch forgetting, exponentially averaged Gram and cross-moment statistics enter each closed-form update, together with proximal stabilization, under-relaxation, weight decay, and iter- ate averaging. The resulting Regularized and Averaged Memory-Regularized (RA-MR) optimizer admits batch-recursive and batch-Tikhonov solvers and preserves derivative-free, memory-efficient, layer-local updates.