The Price of Locality: Why Forward-Forward Underperforms Backpropagation?
Abstract
The Forward-Forward Algorithm (FFA) replaces backpropagation (BP) with layer-wise local contrastive objectives, eliminating the backward pass, yet suffers a persistent performance gap with BP that worsens with depth. This paper diagnoses two distinct deficits: an irreducible optimization floor arising from concurrent local updates, whose magnitude is amplified when kernel contraction degrades the optimization Gram; and a geometric collapse of layer representations driven by kernel contraction itself. On the optimization side, we prove that the FFA loss satisfies the Polyak--{\L}ojasiewicz inequality at each layer, but concurrent layer updates create an irreducible error floor that grows with depth and is amplified when the Gram degrades. On the representational side, the pairwise similarity kernel of layer representations contracts exponentially toward rank one as depth increases, collapsing the diversity of per-layer error signals. This collapse bounds FFA's \emph{effective learning capacity}---the total diversity of gradient information across layers---to grow only linearly with depth regardless of width, whereas BP's chain-rule signal preserves per-layer diversity, yielding a capacity that scales with both depth and width.