Tree-grams : Hierarchical $n$-gram models as controlled approximators of language
Aritra Das ⋅ Francesco Cagnetta ⋅ Darshil Doshi ⋅ Victor V. Albert ⋅ Maissam Barkeshli ⋅ Andrey Gromov
Abstract
We introduce tree-grams, trainable $n$-gram models that factorize the conditional probability tensor $p(x_n \mid x_{1:n-1})$ through a hierarchy of rank-2 and rank-3 merges. A single rank hyper-parameter, $d$, controls model complexity, yielding an interpretable and systematically scalable family of approximators of language. Across all tested orders, tree-grams outperform existing $n$-gram baselines. They also converge to $n$-gram entropies obtained from deep neural network-based language models upto $n \approx 8$, and are competitive at longer context. We ablate the role of different components in the model and demonstrate the emergence of sparse and dense sub-structures. In addition, when allowed to route across candidate topologies, tree-grams exhibit an emergent preference for sparse solutions, with a broken power-law distribution characterizing the path ranks. Finally, we show that tree-grams are useful as controlled synthetic data: tuning the peak learning rate on small proxies trained across a range of tree-gram complexities and transferring the median recovers the optimum of 15x and 20x wider models under $\mu$P.
Chat is not available.
Successful Page Load