VENUS: Decoupling the Samples in MARS Benefits Small-Batch Training
Abstract
AdamW has long been a standard optimizer for neural network training, while recent alternatives include Muon, which uses matrix-aware updates, and MARS, which incorporates a scaled variance-reduction correction. We interpret the unclipped MARS estimator as a convex combination of momentum SGD and momentum variance reduction. When both branches use the same minibatch, independently varying their mixing weights does not expand the estimator family. We introduce VENUS, which generalizes MARS by evaluating the two branches on independent minibatches. Its additional sample-splitting parameter expands the estimator family and recovers MARS as a special case. Combined with spectral-norm updates of the Muon family, VENUS integrates sample-decoupled variance reduction with matrix-aware optimization. We establish convergence guarantees for the broader class of updates defined by a linear minimization oracle (LMO), with an explicit variance factor minimized by balanced weighting. Experiments on GPT-2-style language model pretraining under matched gradient budgets support balanced weighting and reveal a batch-dependent comparison: exact VENUS achieves lower final validation loss at the smallest tested batch size, the difference remains inconclusive at an intermediate batch size, and exact MARS performs better at the largest. Holding the training state fixed, we separate what averaging over independent gradients contributes from what the covariance between the refresh and transport terms contributes, and show how that covariance moves the split that minimizes variance. A second comparison, at equal computational work, shows that lower variance is not the same as faster progress: which estimator comes out ahead depends on whether the update is Euclidean or comes from a layerwise oracle. The observed small-batch benefit motivates further study of training under limited computational resources.