TorchTangents: An Accessible Bridge for Architecture-Preconditioned Optimizers
Abstract
Standard gradient descent operates natively in parameter space, remaining mathematically blind to a neural-network's particular structure. Recent optimization strategies, such as Muon, alter this dynamic by enforcing layer-wise spectral constraints to enable the early amplification of weak signals that would otherwise be dominated by batch noise. In this paper, we further posit that explicit architecture-preconditioned optimization is an intriguing frontier for stabilizing wide minima. However, there exists currently a gap between how to seamlessly extract meaningful architectural biases useful for preconditioning training. We propose NTK-related architectural properties could serve as a promising family of signals to utilize as a training preconditioner. To achieve this, we introduce TorchTangents, a PyTorch-native engine designed to bridge infinite-width neural tangent kernel theory with finite-width empirical optimization. The library introduces an API to dynamically convert standard network layers into their mathematically exact infinite-width kernels without requiring architectural rewrites, and in turn, serves as a new foundation upon which architecturally-preconditioned optimizers can be built from. We empirically demonstrate with preliminary results that a novel optimizer, Architecture-Aware-Orthogonal-Descent (AAOD), yields greater test-time model performance than when trained with Adam or Muon alone. TorchTangents is available at: https://anonymous.4open.science/r/torchtangents-B747/README.md