SiliconBench: Speed, Memory, and Fidelity for LLM Inference on Apple Silicon
Ranran H Zhang ⋅ Aysa Fan ⋅ David M Correia ⋅ Alex Cheema ⋅ Rui Zhang
Abstract
Apple Silicon is a major consumer platform for local LLM inference, with at least ten serving engines competing for users. Speed-only leaderboards on this platform mislead: the fastest stack often claims memory the rest of the machine needs, or silently produces wrong output. We introduce \ourbench, a benchmark that evaluates ten Apple Silicon inference stacks through three lenses: throughput and latency at concurrency 1, 8, and 16; peak memory under contention with the operating system; and fidelity as weighted F1 on a classification task against an NVIDIA reference run on identical weights. The benchmark covers three model releases and two workload types (short chat and multi-turn agent prompts averaging ${\sim}$4K input tokens). A maintainer agent re-runs the full benchmark weekly; in its first five runs it caught a single upstream library bump that broke three of ten stacks through three distinct failure modes and diagnosed a chat-template error that corrupted one stack's fidelity scores. No audited stack performs well across all three lenses, and the stack rankings shift when memory and fidelity are read alongside speed. The candidate set, stacks without disqualifying gaps under any lens, reduces from ten to four. A separate multi-node comparison shows that tensor parallelism over RDMA scales decode $1.3\times$ on two nodes, while pipeline parallelism over TCP regresses $16$--$19\%$. In every case the gaps are in the engines, not the hardware: Apple Silicon provides sufficient memory, compute, and interconnect; the engines have not yet exploited it. We release the harness, per-run results, and weekly journals.
Chat is not available.
Successful Page Load