arXiv:2609.33385v1 Announce Type: new Abstract: Semi-autoregressive diffusion large language models (dLLMs) improve decoding parallelism through iterative block-wise denoising, but scaling them with mixture-of-experts (MoE) layers introduces a large expert parameter footprint that exceeds memory-constrained GPU capacity. Expert offloading is a natural remedy, yet existing MoE serving systems targ
OLED-MoE is a newly proposed expert offloading system designed to accelerate Mixture-of-Experts (MoE)-based diffusion Large Language Models (dLLMs) by leveraging inter-iteration expert locality. Unlike traditional systems that rely on intra-iteration prefetching, OLED-MoE shifts the optimization target to expert retention across denoising iterations, utilizing token confidence signals to predict and preserve high-value experts in GPU memory.
Key performance improvements include: 1.23×–7.93× reduction in time per output token (TPOT). 1.44×–4.23× improvement in expert cache utilization. * Achieving performance close to full-residency models while using only 40% of the expert GPU memory.
The system addresses the bottleneck where semi-autoregressive dLLMs activate up to 8× more experts per iteration compared to autoregressive models, making standard prefetching ineffective. By exploiting the high overlap (85–90%) of activated experts between adjacent denoising steps, OLED-MoE significantly reduces CPU-GPU transfer overhead. The paper, submitted on September 27, 2026, has been accepted by EuroSys '27.
OLED-MoE targets a systems bottleneck that appears when mixture-of-experts (MoE) layers are added to diffusion large language models (dLLMs). dLLMs use semi-autoregressive, iterative block-wise denoising to generate multiple tokens in parallel, which can improve decoding throughput compared with strictly autoregressive generation. However, MoE scaling dramatically increases the expert parameter footprint, often exceeding the memory capacity of a single GPU or memory-constrained accelerator. The paper argues that this problem is not adequately addressed by existing MoE serving or expert-offloading systems, which are typically designed around autoregressive token generation or generic expert activation patterns rather than the repeated, structured denoising iterations used by dLLMs.
The central contribution is an inter-iteration locality-aware expert offloading framework. Rather than treating each denoising step as an independent MoE forward pass, OLED-MoE exploits the fact that expert usage can exhibit strong temporal and structural locality across iterations: experts activated in one denoising step are often relevant again in nearby steps or related blocks. The system uses this locality to decide which experts should remain resident in GPU memory, which should be prefetched from host memory, and which can be evicted. In doing so, it aligns expert placement and data movement with the dLLM inference pipeline, reducing unnecessary host-GPU transfers and improving the overlap between expert loading and model computation.
This matters because it addresses a practical barrier to deploying large MoE-based dLLMs on realistic hardware. Naive offloading can be dominated by memory traffic, while autoregressive MoE serving systems may miss the temporal reuse patterns that are central to iterative denoising. By making expert offloading aware of dLLM-specific execution behavior, OLED-MoE enables larger MoE models to run within tighter GPU memory budgets and potentially with lower inference latency. More broadly, the work highlights a broader systems insight: as non-autoregressive and diffusion-style language models become more common, inference engines must be redesigned around their unique activation and memory-access patterns rather than simply reusing techniques developed for autoregressive LLMs.