Native Multi-Dimensional Subquadratic Operators via Input Dependent Long Convolutions
David Wessels ⋅ FARHAD RAMEZANGHORBANI ⋅ David W Romero ⋅ Alireza Moradzadeh ⋅ Olivia Viessmann ⋅ Maksim Zhdanov ⋅ John St. John ⋅ Ken Janik ⋅ David Knigge ⋅ Yucheng Tang ⋅ Erik Bekkers ⋅ Saee Paliwal
Abstract
Subquadratic alternatives to attention require compromises when applied to multi-dimensional data: standard convolutions abandon the global, input-dependent receptive field, while recurrent models require rasterizing images, volumes, and PDE grids into an ad-hoc $1\rm D$ scan order that violates their spatial structure. We introduce HyenaND, an operator that acts directly on native $\rm ND$ data in its intrinsic geometry and recovers a global, input-dependent receptive field at subquadratic cost. It uses an implicit, input-dependent parameterization of the multi-dimensional convolutional kernel. We provide a CUDA implementation, \texttt{nSubQ}, which fuses the FFT-convolution path to turn HyenaND's $O(N \log N)$ scaling into wall-clock speedups. Across long-context genomics, computer vision, medical imaging, and PDE modeling, pure HyenaND stacks match the accuracy of strong attention-based baselines, while hybrid configurations that interleave HyenaND and attention layers outperform both pure attention and strong recurrence-based hybrids.
Chat is not available.
Successful Page Load