Parallelising Matrix-Products in Isotropic Multilayer Perceptrons
Abstract
Isotropic activation functions are characterised by a constraint to standard orthogonal-group equivariance. This unusual property enables this form of activation function to have a scalar non-linear term that commutes with weights. This allows the typical sequential series of weights and biases to be contracted into single matrices and vectors upfront and in parallel, whilst retaining a non-linear scalar term sequentially. Thus, the possibility of precomputing all matrix-matrix and matrix-vector operations in parallel is made clear. These networks are characterised by a deep nonlinear dependence whilst not requiring a deep matrix-multiplication formulation. Furthermore, the MLP architecture is expanded to a broader functional class, enabling this speed-up at both training and inference time, with a performance advantage demonstrated in more complex image-classification tasks.