AdaPrecise: A Task-Agnostic Dynamic Precision Routing Framework via Gumbel-Softmax for Edge Inference
Abstract
The growing size of deep neural networks (DNNs) makes deployment on resource-constrained edge devices difficult. Static quantization is the standard remedy for memory and compute bottlenecks, but it implicitly assumes that all inputs and intermediate features have similar representational complexity. As a result, static methods over-allocate precision to easy background features and under-allocate it for the rare patterns that actually drive task performance, yielding a sub-optimal accuracy–latency trade-off. Recent dynamic-computation methods address part of this problem through token routing or channel gating, but they are typically tied to a single modality and rely on theoretical Floating-Point Operations (FLOPs) as the efficiency metric—a proxy that often disagrees with measured edge latency because of memory unpacking overhead on commodity hardware. We propose AdaPrecise, a task-agnostic, instance-aware framework for dynamic precision routing. By representing heterogeneous DNN architectures as a unified Directed Acyclic Graph (DAG), AdaPrecise inserts a lightweight, modality-independent router that assigns a bit-width to each computational node at runtime. Discrete routing is made differentiable through an annealed Gumbel-Softmax estimator. To close the gap between algorithmic cost and real latency, we introduce a hardware-aware Look-Up Table (LUT) loss that penalizes the measured execution time of each precision on the target device, suppressing low-bit choices that are cheap in theory but slow in practice. Combined with a multi-stage knowledge-distillation curriculum that prevents router mode collapse, AdaPrecise improves the accuracy–latency trade-off across vision, NLP, and audio benchmarks, supporting the claim that the framework is broadly applicable rather than modality-specific.