The Device Decides: Benchmarking the Reliability of Agentic SLMs at the Edge
Jiajun Wu ⋅ Jiayu Zhou ⋅ Steve Drew
Abstract
Small language model (SLM) agents now run directly on edge devices, where the model to deploy is usually chosen by capability. Capability, however, does not tell the full story. After deployment, the latency tail, output stability, and calibration depend as much on the hardware and software stack as on the model. We call these on-device properties reliability. To our knowledge, no capability benchmark reports reliability. We present LegitOnEdge, a novel benchmark that measures the reliability of an agentic SLM on the edge device where it will run. We use it to evaluate five instruction-tuned SLMs of 3.8-8B parameters on an NVIDIA Jetson Orin Nano and a DGX Spark across four workloads. We show that the same model completes $0.54$ of a multi-step agentic task suite on one device but $0.04$ on the other. In addition, the model outputs are only $0$-$26\%$ byte-identical across the two builds. Across twenty cells, our reliability score over latency, throughput, and energy shows no detectable association with capability, and its confidence interval rules out a strong association in either direction. We argue that an agentic SLM deployment should therefore be ranked by reliability of the model measured on its target device, not by the model alone.
Chat is not available.
Successful Page Load