Modern reinforcement learning (RL) post-training pipelines for large language models (LLMs) increasingly combine rollout workloads across multiple domains and feedback paradigms. Prefix-aware routing improves inference efficiency through cache reuse and load balancing, but it does not control how heterogeneous rollout sessions compete for KV-cache capacity. When reinforcement learning with verifia

Topological visualization of Scheduling Mixed RL Rollouts Beyond Prefix Locality
Brave API

Modern RL post-training systems address the inefficiencies of mixed-domain rollouts by moving beyond static prefix locality to dynamic, trajectory-centric and length-aware scheduling. Heddle utilizes a trajectory-centric design with progressive priority scheduling to minimize queueing delays for long-tail trajectories, while RollPacker employs tail batching to consolidate prompts leading to long responses into designated steps, reducing GPU idle time.

Other frameworks optimize this further through pipeline multiplexing and asynchronous overlap. JigsawRL exposes intra-stage imbalance via a Sub-Stage Graph to dynamically allocate resources and migrate long-tail rollouts, whereas RolloutPipe implements complete-group pipelining to start training before rollout completion. Additionally, SortedRL uses online length-aware scheduling to reorder samples by output length, creating micro-curricula that maintain on-policy stability while boosting throughput.

Generated Aug 22, 2026
Open-Weights Reasoning

The paper addresses a systems bottleneck in large-scale RL post-training for LLMs: modern pipelines increasingly run heterogeneous rollout workloads—spanning different domains, prompt distributions, response lengths, and feedback or verification mechanisms—on shared inference infrastructure. Prefix-aware routing can improve efficiency by exploiting cache locality and balancing load across replicas, but the authors argue that this is insufficient when many rollout sessions compete for a finite KV-cache. Prefix locality captures whether a request can reuse cached state, but it does not determine whether admitting that request will evict more valuable state, stall other sessions, or degrade overall rollout throughput.

Its central contribution is to reframe rollout scheduling as a KV-cache-aware resource-management problem rather than merely a prefix-matching or load-balancing problem. The key insight is that mixed RL rollouts should be scheduled according to how they affect shared cache capacity: which prefixes are worth retaining, which sessions are likely to produce long or short continuations, how different feedback paradigms create variable downstream latency, and how admission decisions interact with eviction and session completion. In effect, the paper pushes scheduling “beyond prefix locality” by treating cache residency, eviction cost, and workload heterogeneity as first-class factors in deciding how to route and admit rollouts.

This matters because RL post-training is becoming less uniform and more pipeline-intensive, with multiple task domains and reward/verification mechanisms running concurrently. Without cache-aware scheduling, prefix-aware systems can still suffer from cache thrashing, uneven replica utilization, straggler rollouts, and reduced effective GPU throughput. By accounting for the competitive dynamics of mixed rollouts, the work offers a more robust foundation for scaling RL training pipelines where inference efficiency and cache management directly determine end-to-end training speed.

Generated Aug 23, 2026
Sources