Vision-language models (VLMs) enable embodied agents to reason and act from visual observations and language instructions. Reinforcement learning (RL) post-training enhances these capabilities using task feedback, but current on-policy RL runtimes execute rollout, reference scoring, and actor training in strict serial phases. While effective for text-only RL, this phase-granular execution is waste

Topological visualization of Rollplex: Cross-Phase GPU Spatial Sharing for Vision Language Model Post-Training
Brave API

Rollplex is a runtime system that accelerates Vision-Language Model (VLM) post-training by spatially overlapping GPU computation across phases that are typically executed serially. It achieves this by moving prefix computation (such as video encoding and prompt prefill) into the rollout decode window, thereby utilizing GPU capacity that would otherwise remain underutilized.

Implemented on 32 H800 GPUs, Rollplex delivers a 1.23×–1.30× speedup over serial colocation and a 1.57×–2.24× speedup over disaggregated setups. These gains are realized through two key mechanisms: phase-aware memory management, which controls High Bandwidth Memory (HBM) residency based on data lifetimes, and parallelism-aware weight sharing, which allows rollout and training engines to use different tensor-parallel degrees while sharing actor weights. This approach preserves synchronous on-policy semantics and autograd correctness without requiring redundant model copies.

Generated Sep 1, 2026
Open-Weights Reasoning

Rollplex addresses a systems bottleneck in RL post-training for vision-language models (VLMs): the fact that on-policy RL pipelines often run rollout generation, reference-model scoring, and actor training as strictly sequential phases. This serial execution can leave GPUs underutilized, especially in VLM settings where workloads are heterogeneous, memory-intensive, and uneven due to variable visual inputs, token lengths, and multimodal batching. The paper’s central idea is to move from phase-granular temporal scheduling to cross-phase GPU spatial sharing, allowing different stages of the RL loop to execute concurrently on partitioned or co-scheduled GPU resources rather than waiting for one phase to finish before the next begins.

The key contribution is a runtime or scheduling framework that overlaps these normally disjoint phases while managing the associated contention, memory pressure, and dependency constraints. By spatially sharing GPUs across rollout, scoring, and training, Rollplex aims to reduce pipeline bubbles, improve device utilization, and shorten end-to-end RL iteration time. The broader insight is that the performance ceiling of VLM RL post-training may be limited less by the RL algorithm itself and more by how the surrounding execution runtime schedules model work across accelerators.

This matters because VLM post-training for embodied agents is increasingly compute-bound and expensive, and faster RL loops can directly affect how quickly models can be improved with task feedback. If Rollplex can deliver meaningful throughput or latency gains, it makes large-scale VLM RL more practical and economically viable, and it also has implications beyond VLMs for any multi-phase ML system that interleaves inference, scoring, and training workloads on shared GPU infrastructure.

Generated Sep 1, 2026
Sources