Sparsely Wired Mortal LLM Inference
Abstract
This paper rethinks Large Language Model (LLM) inference from a Mortal Computing perspective, in which computation is embodied in hardware and functionality is inseparable from its physical substrate. Rather than treating inference optimization as merely a matter of arithmetic simplification or precision reduction, we explore hardwiring as a way to eliminate weight access at the root of data movement. We show that viable hardwiring begins not by retaining weights in their multiplicative form, but by factorizing them into combinations of shiftable bases and rewiring computation at the base level, where redundancy is exposed. This view, however, introduces a fundamental challenge in the dramatic growth of wires required to reconstruct weights. To address this, we propose Moira, a budgeted greedy search that formulates wire-count control as a Lagrangian relaxation problem and selects only the bases that satisfy the objective under a wiring budget. What sets Moira apart is its adaptive base selection, which gives it a dual characteristic of pruning and quantization while preserving model expressivity under sparse wiring without retraining. Extensive evaluations demonstrate that Moira achieves near-lossless performance with, on average, only about two bases per weight and outperforms existing compression methods. Moira opens a concrete path beyond the von Neumann paradigm, marking a first step toward Digital Mortal Computing for LLMs.