ElasticInfer: Safe Contextual Model Routing under Dynamic Edge-Resource Constraints
Abstract
On-device intelligence for embodied agents must operate under highly dynamic resource constraints caused by shifting concurrent workloads and thermal throttling. Static deployment strategies either risk severe memory and latency violations or conservatively sacrifice task accuracy (such as perception quality). We present ElasticInfer, an adaptive framework that decouples resource admission from model selection to maximize perception accuracy under fluid system constraints. Its Nano-first policy quickly extracts reusable detections and input context, then utilizes a contextual Gaussian Thompson sampling strategy to select the optimal model from admitted detector tiers. We evaluated ElasticInfer on 5,000 COCO validation images under simulated resource dynamics. Our approach (Elastic-TS) achieves 0.406 mAP at a modeled mean latency of 12.3 ms, significantly outperforming a conservative baseline cascade which yields 0.372 mAP at 3.6 ms. Operating with independently calibrated synthetic cost margins, guarded Elastic-TS records zero modeled latency violations, while successfully delivering 97.05\% of queries. Finally, in a physical GPU evaluation featuring concurrent compute, managed allocation, and warmed input shapes, ElasticInfer successfully delivers 100\% of its 180 queries within a strict 50 ms deadline, experiencing no failures and a low mean latency of 21.31 ms. These results demonstrate the critical value of explicit admission and delivery accounting, while comparisons with baseline routers further characterize the remaining accuracy–cost trade-off.