SABLE-X: A Sparsity-Preserving GPU System for Batched Differentiable Nonlinear Solves in End-to-End Learning
Suho Park ⋅ Keunju Song ⋅ Hongseok Kim
Abstract
End-to-end learning with implicit layers for equality constraints repeatedly solves batched sparse nonlinear systems and differentiates through their solutions, making runtime and memory critical. In this regard, we present SABLE-X, a PyTorch system for batched sparse nonlinear solves with implicit differentiation. From local residual templates and sparse connectivity, it generates problem-specific GPU kernels and reusable block-diagonal sparse layouts. The layouts and solver analyses for the forward and transpose systems are reused across repeated solves. Across six benchmarks, SABLE-X achieves up to $575\times$ the forward--backward throughput of the fastest converged baseline at the same batch size. On the two largest systems, it scales to a batch size of $1{,}024$, whereas several baselines exhaust memory or fail to converge. Across three end-to-end learning pipelines, replacing the baseline solver layers with SABLE-X achieves $517\times$ to $4{,}086\times$ the baseline throughput and supports training batch sizes $32\times$ to $64\times$ those of the corresponding baselines.
Chat is not available.
Successful Page Load