Evaluating the Applicability of LLM Inference Optimizations on Apple Silicon
Kyrylo Yemets ⋅ Andrii Yaroshevych ⋅ Victor Muryn ⋅ Ostap Khomenko ⋅ Pavlo Kryven ⋅ Maksym Kmet ⋅ Maksym Shamrai ⋅ Mariya Hirna
Abstract
Apple Silicon is among the most widely used platforms for local LLM inference, yet nearly all inference optimizations are validated on server-class NVIDIA GPUs under batched serving workloads. Which optimizations retain their benefits when moved to Apple Silicon, and what determines their applicability? We evaluate four optimization families: KV-cache pruning, speculative decoding, ported GPU kernels, and post-training quantization. Experiments span four Apple machines under a common single-request decode workload, with matched RTX~4090 comparisons where possible. KV-cache pruning can reach up to $1.75\times$ end-to-end throughput on Apple but can reduce throughput on desktop NVIDIA GPUs. In matched experiments, all evaluated speculative decoding methods reduce throughput on Apple, despite reported $2$-$6\times$ gains on server GPUs. CUDA kernel designs transfer only partially: public Metal lacks asynchronous copy, and batch-one decode makes several serving-oriented mechanisms inapplicable, although transferable components yield up to $1.20\times$ decode throughput. Quantization gains vary by machine. Across four families, a reported speedup is not a portable property of a technique alone but a joint function of technique, runtime, hardware contract, and workload regime.
Chat is not available.
Successful Page Load