The Gap Is the Loop: What Background Work Really Costs On-Device LLM Decode on Apple Silicon
Abstract
On-device inference on Apple Silicon shares a singular DRAM pool and power budget. We measure what exactly that costs a batch-1 8B decoder on a base M4 under three synthetic loads that isolate memory bandwidth, CPU, and GPU (stream, fma, gpualu) and eight real background applications, across eleven pre-registered follow-up experiments. Bandwidth contention, as expected, slows decode by 72% when saturated. The relationship between memory pressure and latency degradation is convex and non-linear so performance worsens as memory saturation approaches 100%. However, we found that 6 out of 8 light background applications make a synchronous per-token decode loop faster, which has been replicated at n=17-20. The cause is that a synchronous loop leaves the GPU idle for about 6% each step while the submitting CPU core downclocks. But light background tasks kept the CPU cores awake at higher clock frequencies and thus hides the delay. Pipelining the loop and making it asynchronous, (submitting the next token (n+1) to the GPU before reading the current token (n), as mlx_lm does) 95% of the idle gap is eliminated and raw generation speed increases by 9 to 12%. The initial speedup now correctly demonstrates a penalty in performance by about +1 to +4 percentage points. Benchmarks that time tokens synchronously thus report background interference as beneficial rather than harmful. We also explore the Quality of Service (QoS) setting. Operating systems use QoS to demote background tasks so they do not limit and interrupt foreground work. Counter-intuitively, it was measured that lowering a background job's QoS makes it more damaging to active LLM decoding, as running the same encode on the same frame rate on E-cores in background slowed decode by +13.9 pp more than foreground P-cores.