Quantifies accuracy degradation from layer dropout in large-scale LLM pre-training and proposes mitigations to restore its benefits.
Don't Drop Dropout: Optimizing Layer Sparsity for Efficient LLM Training and Inference challenges the recent trend of removing layer dropout from LLM pre-training by demonstrating that it can be effectively utilized with proper optimization. The study reveals that while layer dropout can degrade accuracy if mismanaged, optimal layer distribution, time schedules, and optimizer hyperparameters allow a 3.9B-parameter LLM to achieve lower validation loss while saving 20% of training FLOPs.
Furthermore, the paper highlights significant post-training benefits, noting that layer dropout enables optimizations like early exit and intermediate-layer skipping, which yield up to a 1.7x inference speedup with negligible accuracy loss. These findings, derived from over 2,400 training experiments spanning models from 271M to 3.9B parameters, establish best practices for integrating layer dropout into state-of-the-art LLM training regimes.
Problem and scope. The paper studies layer dropout as a form of dynamic architectural sparsity for large language models, where selected transformer layers are skipped during training or inference to reduce FLOPs and improve throughput. It frames this as a central efficiency–quality trade-off: deeper LLMs are costly to train and serve, but naively reducing effective depth can damage model capability. The work quantifies the accuracy degradation caused by layer dropout in large-scale pre-training and examines how the penalty depends on factors such as model scale, skip rate, layer placement, and training dynamics.
Key contributions. Its main contribution is a more rigorous account of what layer sparsity buys and what it costs. Rather than treating layer skipping as a simple compute-saving trick, the paper characterizes the mechanisms behind its accuracy loss—such as reduced effective capacity, disrupted residual computation, and instability in the training process—and proposes mitigations that preserve the efficiency benefits while restoring performance. The resulting approach is aimed at optimizing which, when, and how aggressively layers can be dropped, rather than merely reducing depth in a static or uncontrolled way.
Why it matters. This matters because efficient LLM training and inference are increasingly constrained by compute cost, latency, and deployment budgets, and dynamic depth is a promising but under-characterized lever. By quantifying the accuracy penalty of layer dropout and offering practical mitigations, the paper gives researchers and engineers a more reliable path toward sparse-depth models that retain the benefits of large-scale pre-training. More broadly, it speaks to a key design question in efficient LLMs: how to make model depth and capacity dynamically adjustable without sacrificing the stability, generalization, and downstream performance expected of modern foundation models.