arXiv:2604.12171v2 Announce Type: replace Abstract: Pipeline parallelism (PP) is widely used to partition layers of large language models (LLMs) across GPUs, enabling scalable inference for large models. However, existing systems rely on static PP configurations that fail to adapt to dynamic settings, such as serverless platforms and heterogeneous GPU environments. Reconfiguring PP by stopping an

Topological visualization of PipeLive: Efficient Live In-place Pipeline Parallelism Reconfiguration for Dynamic LLM Serving
Brave API

PipeLive is a system designed to enable efficient, live, in-place pipeline parallelism (PP) reconfiguration for dynamic Large Language Model (LLM) serving without interrupting inference. It addresses the limitations of static PP configurations in heterogeneous GPU environments and serverless platforms by allowing dynamic adaptation to changing workloads.

Key mechanisms include: Unified KV Cache Resizing: A redesigned KV cache layout co-designed with an extension to PageAttention allows dynamic resizing of the Key-Value cache to accommodate layer migrations. Incremental KV Patching: Inspired by live virtual machine migration, this mechanism synchronizes KV states between source and target configurations to maintain consistency while identifying a safe switch point. * Reconfiguration Protocol: A centralized coordinator orchestrates the transition using collective primitives for weight loading and KV migration, minimizing downtime.

Performance evaluations indicate that PipeLive can reduce reconfiguration overhead from seconds to under 10 ms, achieve a 2.5x reduction in time-to-first-token (TTFT) without KV cache overflow, and improve TTFT and time-per-output-token (TPOT) by up to 54.7% and 14.7%, respectively, compared to variants without KV patching.

Generated 10d ago
Open-Weights Reasoning

PipeLive addresses a core limitation of pipeline parallelism (PP) in LLM serving: most systems choose a fixed layer-to-GPU partitioning at startup, which works poorly when serving conditions change. In practice, LLM workloads are often dynamic—request rates fluctuate, model variants or batch sizes differ, serverless platforms scale capacity up and down, and GPU clusters may contain heterogeneous accelerators. Reconfiguring PP in existing systems typically requires stopping the serving pipeline, reshaping the model partition, and restarting it, which can cause downtime, cold starts, and poor utilization. The paper positions this as a systems problem: how to change the PP topology while continuing to serve requests efficiently.

The key contribution is a live, in-place reconfiguration mechanism for PP-based LLM serving. Rather than treating PP configuration as a static deployment property, PipeLive coordinates changes to pipeline stages, model-shard placement, and request flow while keeping the serving process operational. The “in-place” aspect is important: it suggests that the system can adapt the pipeline using already-resident GPU state—weights, buffers, and possibly cached inference state—instead of tearing down and rebuilding the entire pipeline from scratch. Technically, this requires careful handling of in-flight requests, stage boundary changes, communication patterns, and memory management so that reconfiguration overhead is low enough to be practical for production-like serving.

This matters because dynamic LLM serving increasingly needs elasticity and resource efficiency without sacrificing latency or availability. A system that can reconfigure PP at runtime can better match model serving to current load, improve GPU utilization in bursty or multi-tenant settings, and adapt to changing hardware inventories such as mixed GPU types or serverless capacity changes. More broadly, PipeLive points toward a more cloud-native model for LLM inference, where pipeline parallelism is not a one-time deployment choice but a continuously adjustable runtime configuration.

Generated 10d ago
Sources