arXiv:2610.01950v1 Announce Type: new Abstract: Sparse expert activation reduces MoE models' computation, yet expert weights can exceed limited device memory. Offloading makes inference feasible on a compact AI appliance but exposes host-to-device transfers to the inference path. We present MoE-CORE, a system that coordinates expert offloading and residency for memory-constrained MoE inference. I
MoE-CORE is a system designed for memory-constrained Mixture-of-Experts (MoE) inference that coordinates expert offloading and residency to mitigate host-to-device transfer latency. It addresses the challenge where expert weights exceed limited device memory by using alternating buffers during the prefill phase and combining nonuniform layer-wise cache capacity, domain-informed initialization, and cross-layer prefetching during decoding.
Key performance metrics on an 84-GB NPU include: DeepSeek-V4-Flash-W4A8: Reduces mean Time Per Output Token (TPOT) from 1268.9–1269.1 ms (vLLM Prefetch) to 38.0–44.8 ms. GLM-5.2-W4A8C8: Reduces mean TPOT from 5941.5–5941.8 ms (vLLM Prefetch) to 206.6–220.5 ms. * Optimization: Under an optional approximate expert substitution with multi-token prediction, MoE-CORE achieves a best-case TPOT of 21.5 ms.
The system distinguishes between prefill and decode phases, reserving separate device-memory regions for each to optimize throughput and latency respectively, while handling low-score cache misses via score-based expert substitution.
MoE-CORE addresses a key deployment bottleneck for mixture-of-experts (MoE) models on memory-constrained accelerators: even though sparse expert activation reduces per-token compute, the total expert weight footprint can still exceed the device’s high-bandwidth memory. Offloading experts to host memory makes such models loadable on compact AI appliances, but if host-to-device transfers are left unmanaged, they become a dominant latency source in the inference path. The paper therefore frames expert offloading and residency as a coordinated systems problem rather than a simple memory-management afterthought.
The central contribution is a system-level approach that jointly manages which experts reside on-device, which are evicted to host memory, and how transfers are scheduled relative to inference execution. The underlying insight is that expert routing in MoE models creates exploitable structure: expert demand is not uniformly random, and residency decisions can be made proactively or adaptively to keep frequently or imminently needed experts available while minimizing exposed transfer time. By coordinating memory placement with the inference pipeline, MoE-CORE aims to reduce the latency penalty of offloading and make large MoE models practical on devices with limited memory.
This matters because MoE architectures offer a promising path to scaling model capability without proportionally increasing compute cost, but their deployment has often been constrained by accelerator memory capacity. A system that can efficiently manage expert residency on compact hardware could enable local, private, or edge inference for larger MoE models, reducing dependence on high-end datacenter accelerators and making sparse-model inference more accessible in memory-limited settings.