arXiv:2610.01265v1 Announce Type: new Abstract: The widespread adoption of Mixture-of-Experts (MoE) has created a growing need for deployment on heterogeneous platforms. However, it exposes a fundamental mismatch between the algorithmic demands of large-scale MoE and the disparate characteristics of hardware.Existing CPU-GPU hybrid inference systems fail to resolve this as they either encounter P

Topological visualization of RapidMoE: Exploiting Cross-Asymmetry via Adaptive Residual Offloading for Large-Scale MoE Inference
Brave API

RapidMoE (arXiv:2610.01265) is a system for efficient large-scale Mixture-of-Experts (MoE) inference that exploits cross-asymmetry between algorithmic routing skew and hardware resource disparity. It shifts offloading from the expert level to the bit level via a RESplit framework, which decomposes weights into memory-intensive residuals (offloaded to CPU) and quantized components (retained on GPU).

Key innovations include: Unified Multi-Level Importance Arbitration (UMIA) for runtime adaptive discrimination of critical experts. Hardware-aware workload partitioning that aligns storage-heavy tasks with CPU capacity and compute-intensive tasks with GPU throughput. * Performance gains of up to 3.5x in decoding and 2.1x in prefilling compared to state-of-the-art offloading systems.

The paper, authored by Wenxun Wang et al., was accepted by EuroSys 2027 and submitted in October 2026.

Generated 13h ago
Open-Weights Reasoning

RapidMoE targets a practical deployment bottleneck for large Mixture-of-Experts models: although MoE architectures improve efficiency by activating only a subset of experts, their full parameter footprint often exceeds the memory of a single accelerator. The paper focuses on heterogeneous CPU–GPU inference, where the CPU provides large, low-cost memory and the GPU provides high compute throughput, but where naïve hybrid systems can suffer from poor expert placement, load imbalance, and transfer latency. It frames this mismatch as a cross-asymmetry problem: the irregular, sparse activation pattern of MoE experts does not align well with static or one-sided offloading strategies that treat CPU and GPU as interchangeable resources.

The key contribution is an adaptive residual offloading mechanism that exploits this asymmetry at runtime. Rather than relying on a fixed partition of experts or a rigid offloading policy, RapidMoE adjusts which expert computations and weights remain on the GPU versus which are offloaded to the CPU, based on activation patterns and system state. The central insight is that MoE’s sparsity makes it possible to keep frequently used experts resident on the GPU while moving less active or “residual” experts to CPU memory, then overlapping or managing transfers to reduce the impact on latency. This turns the CPU–GPU hardware mismatch from a liability into a design opportunity, aiming to improve GPU memory utilization, effective throughput, and tail latency for large-scale MoE inference.

The work matters because large MoE models are increasingly important for serving high-capability language and multimodal systems, but their memory requirements make deployment on GPU-only systems expensive or infeasible. By focusing on CPU–GPU hybrid inference, RapidMoE addresses a realistic cost-constrained setting where heterogeneous nodes can host models that would otherwise require multiple high-end accelerators. If the system achieves meaningful gains in latency or throughput, it provides both a practical serving optimization and a broader systems design lesson: for sparse expert models, deployment performance depends not only on model architecture, but on how well the execution engine exploits the asymmetric properties of available hardware.

Generated 13h ago
Sources