Critical Points of Linear Self-Attention with Quadratic Loss on a Single Training Example
Ava Berenji ⋅ Khang Nguyen ⋅ Guido Montufar
Abstract
We consider the parameter optimization problem for a linear self-attention mechanism trained on a single training example under quadratic loss. In the single-head, single-layer setting, the model is cubic in the input tokens and multilinear in the query, key, and value parameter matrices. We provide a complete characterization of the critical points of the optimization problem. We first reduce the problem to the critical-point analysis of a deep linear network, for which a complete characterization is known, and then lift the resulting critical points back to the original parametrization through a reduction map based on the singular value decomposition of the input matrix.
Chat is not available.
Successful Page Load