Cheap Online Monitoring Helps: Case Study on nanochat
Maksym Khavil ⋅ Ryan Y Xu ⋅ Martin Schmid
Abstract
Obtaining the best results from neural network training requires many hours of experimentation with different configurations. Much of this experimentation is guided only by loss curves and task metrics such as accuracy, with little insight into the model's failures. Many modern domains also require task-specific architectures, creating additional obstacles to stable convergence. In this paper, we discuss a set of metrics for introspecting model training and implement them in a plug-and-play PyTorch library, OurLibrary. We use these metrics to identify an anomaly in the value-embedding pathway of Karpathy's nanochat pretraining harness. Guided by these findings, we propose an intervention that yields a $1.6\%$ relative improvement in validation bits-per-byte, $28 \times$ larger than seed noise. We then analyse the effect of the intervention by comparing metrics collected during training.
Chat is not available.
Successful Page Load