Accelerating Tensor Products in E(3)-Equivariant Graph Neural Networks
Manasvi Goyal
Abstract
Accurate atomistic simulations are essential for understanding molecular and materials systems, but first-principles methods are computationally expensive at the spatial and temporal scales required for molecular dynamics. E(3)-equivariant graph neural networks such as NequIP offer a promising alternative by learning interatomic interactions while respecting physical symmetries, such that rotating or translating an atomic system produces corresponding transformations of the model predictions. However, the tensor operations required to preserve these symmetries introduce substantial computational overhead, limiting the efficiency of equivariant models at larger simulation scales. We analyze the computational structure of E(3)-equivariant graph convolutions in NequIP and investigate implementations of their dominant tensor-product computation. Hierarchical profiling across four NequIP-OAM architectures identifies the equivariant interaction layers as the dominant model component and further localizes the primary computational cost to tensor-product operations within the equivariant convolution. Motivated by this analysis, we formulate the underlying tensor-product computation as a standalone kernel and implement it using multiple approaches, including C++/CUDA, JIT compilation, and JAX. In our NequIP setup, e3nn serves as the default backend for equivariant operations, making it a natural reference for evaluating alternative implementations. We therefore benchmark our implementations against e3nn and two existing optimized equivariant libraries, OpenEquivariance and NVIDIA CuEquivariance, using identical model configurations to isolate differences in implementation efficiency. We benchmark these implementations using NequIP-OAM-S/M/L/XL workloads spanning increasing model complexity. GPU experiments are performed on NVIDIA A100 GPUs using a molecular graph containing 1,536 atoms and 46,080 directed edges. We evaluate both forward inference and full forward--backward execution on CPU and GPU, using e3nn as the reference implementation. The benchmarks reveal substantial differences across implementation strategies and hardware backends. Among the approaches implemented in this work, JIT achieves the strongest geometric-mean performance, with a 5.78$\times$ speedup over e3nn on CPU and 7.59$\times$ on GPU across the evaluated configurations, reaching 8.82$\times$ for forward execution on GPU. JIT also exceeds NVIDIA CuEquivariance in our aggregate GPU benchmark while approaching OpenEquivariance, the highest-performing external GPU implementation in our evaluation. This demonstrates that a JIT-based approach can achieve competitive performance without requiring a fully specialized equivariant kernel implementation. Our results connect model-level bottleneck analysis with kernel-level optimization and demonstrate substantial opportunities to reduce the computational cost of E(3)-equivariant graph neural networks. This work provides a path toward more scalable neural interatomic potentials and larger molecular dynamics simulations while preserving the underlying equivariant model and the physical symmetries it encodes.
Chat is not available.
Successful Page Load