Triplet-Aware Sparse Higher-Order Attention for Natural Language Understanding
Abstract
The attention mechanism has managed to have an extreme amount of success since its invention as part of the transformer architecture. This architectural component has been a key driver of the large-scale success beyond previous deep learning architectures in the field of natural language. Despite this widespread success, higher-order generalizations of this mechanism have been rarely explored in the literature. In this work, we extend the attention mechanism to a higher-order querying across tokens triples and test this three-dimensional attention on five classification tasks on natural language data. We further analyze the effect of unstructured and structured sparsity on the performance and stability of 3D attention. We find that with sufficient samples 3D attention consistently outperforms classical 2D attention, opening new directions for further research exploration.