Training Mixture-of-Experts (MoE) models for reinforcement learning (RL) couples two load-balancing problems: sequence composition determines dense attention work in each data-parallel microbatch, while token routing determines sparse expert work on expert-parallel ranks. Optimizing either alone can shift the bottleneck to the other. In MoE RL, rollout-time routing replay exposes every sample's se

Topological visualization of RoutePack: Expert Placement and Attention-Aware Data Packing for MoE Reinforcement Learning
Brave API

Based on the provided search context, there is no mention of a system or method named "RoutePack."

The search results highlight other specific frameworks for addressing load balancing and routing in Mixture-of-Experts (MoE) reinforcement learning training:

  • ReLibra: Uses routing-replay-guided load balancing to handle fluctuating expert demand at micro-batch granularity by placing expert reordering at inter-batch timescales and expert replication at intra-batch timescales.
  • LAER-MoE: Focuses on load-adaptive expert re-layout using a Fully Sharded Expert Parallel (FSEP) paradigm to dynamically reconfigure expert placement during training.
  • MoETuner: Optimizes MoE serving with balanced expert placement and token routing.
  • Towards Stable and Effective Reinforcement Learning for Mixture-of-Experts: Proposes a rescaling strategy guided by router logits to reduce gradient variance and improve convergence stability.

If "RoutePack" is a specific recent or niche work not included in the provided context, I cannot provide details on its mechanisms for attention-aware data packing or expert placement.

Generated Aug 26, 2026
Open-Weights Reasoning

RoutePack targets a load-balancing bottleneck that is specific to Mixture-of-Experts (MoE) reinforcement learning: the training system must simultaneously manage dense attention work across data-parallel microbatches and sparse expert work across expert-parallel ranks. In MoE RL, these two workloads are coupled. A data-parallel batch that looks balanced by token count may still be imbalanced in attention cost because sequences vary in length and attention patterns. At the same time, token routing can concentrate expert activations on a subset of expert-parallel ranks, creating stragglers even when the dense computation is well balanced. The paper’s central insight is that optimizing only one side of this tradeoff can simply move the bottleneck to the other.

The key contribution is a co-design of expert placement and attention-aware data packing. Because RL rollouts can be replayed with routing information, the system can observe each sample’s expert-routing profile before or during training. RoutePack uses this routing replay signal to estimate both the attention work induced by a sequence and the expert load it will place on different expert ranks. It then packs sequences into microbatches and assigns experts to ranks so that the dense attention workload and sparse expert workload are jointly balanced, rather than treating them as independent scheduling problems.

This matters because MoE RL training is increasingly sensitive to tail latency and uneven utilization: long variable-length sequences, nonstationary routing, and large expert-parallel topologies can all amplify imbalance. By exploiting rollout-time routing information and jointly optimizing placement and packing, RoutePack aims to reduce stragglers, improve GPU utilization, and make MoE RL training scale more efficiently. More broadly, the work highlights that modern RL systems for sparse models need scheduling and placement decisions that are aware of both the attention structure of sequences and the routing behavior of expert layers.

Generated Aug 26, 2026
Sources