arXiv:2609.19169v1 Announce Type: cross Abstract: Concurrent local LLM serving on unified-memory desktops must preserve memory headroom and output fidelity, which speed-only rankings overlook. We introduce SiliconBench, which evaluates nine Apple Silicon serving engines through three lenses: speed, memory, and fidelity. We evaluate chat and agent serving on Qwen3, Qwen3.5, and Gemma 4. We use a c
SiliconBench is a benchmark introduced in arXiv:2609.19169 (September 12, 2026) that evaluates local LLM serving on Apple Silicon desktops through three critical lenses: speed, memory headroom, and output fidelity. Unlike traditional metrics that prioritize throughput, SiliconBench addresses the unique constraints of unified-memory systems where inference engines must share RAM with the OS and foreground applications.
The study assesses nine Apple Silicon serving engines (including vLLM, SGLang, llama.cpp, and Ollama) using models like Qwen3, Qwen3.5, and Gemma 4 across chat and agent workloads. Key findings include: Memory Discipline is Critical: Explicit memory budgets do not guarantee headroom; some engines complete requests but approach physical RAM capacity, causing throughput declines. Fidelity Matters: Only three stacks (llama.cpp, vllm-metal, and omlx) satisfied all gates for completion, fidelity, and model coverage against an NVIDIA A100 reference. Concurrency Scaling: Speed-only rankings are misleading; engines like vllm-metal showed superior concurrency scaling (doubling throughput from concurrency 1 to 16), while others like mlx_lm collapsed under agent load. Multi-Node Scaling: Tensor parallelism over Thunderbolt RDMA scaled effectively, whereas pipeline parallelism over TCP regressed in two-machine configurations.
The benchmark proposes three desiderata for Apple Silicon serving: serving architecture readiness (support for new attention ops), memory discipline (preserving headroom), and multi-node scaling. The project includes code and a maintenance workflow for continued validation on GitHub.
SiliconBench is a benchmark for evaluating local LLM serving on Apple Silicon unified-memory desktops, where inference engines must balance throughput, memory pressure, and output quality under a shared memory pool. The material focuses on a gap in existing evaluations: many serving comparisons emphasize speed metrics such as latency or tokens per second, but on desktop hardware those rankings can obscure whether an engine leaves enough memory headroom for concurrent requests or whether its configuration preserves the fidelity of model outputs. SiliconBench addresses this by comparing nine Apple Silicon serving engines across three coupled dimensions—speed, memory, and fidelity—using chat and agent-style workloads on models including Qwen3, Qwen3.5, and Gemma 4.
The key contribution is a more complete evaluation lens for local inference systems. Rather than treating serving performance as a single-axis throughput problem, the benchmark frames desktop LLM serving as a resource-constrained systems problem in which memory residency, concurrency, and output behavior all matter. This is especially relevant for unified-memory architectures, where model weights, runtime state, and host workloads compete for the same memory space, so a fast engine that consumes excessive memory or degrades quality under load may be less useful in practice.
The work matters because local LLM serving on desktops is increasingly important for privacy-sensitive, low-latency, and cost-conscious deployments, but Apple Silicon systems are not simply smaller versions of cloud GPU servers. By measuring speed, memory, and fidelity together, SiliconBench provides a more actionable basis for choosing, tuning, or designing serving engines for unified-memory desktops, and it highlights the tradeoffs that speed-only benchmarks tend to hide.