Long-context LLM applications such as retrieval-augmented generation (RAG) and agentic systems often process tens of thousands of input tokens to produce short outputs, making end-to-end request latency an important serving objective. We show that the maximum number of batched tokens (MBT), which controls the token scheduling budget in vLLM, has a scheduling-pressure-dependent effect on latency. L

Topological visualization of P-PAS: Prefill-Pressure Adaptive Scheduling for Long-Context LLM Serving
Brave API

The search results do not contain information on a system named "P-PAS" (Prefill-Pressure Adaptive Scheduling).

However, the search context highlights several related systems addressing long-context and heterogeneous workload challenges in LLM serving:

  • LAPS (Length-Aware-Prefill Scheduling) A system that disaggregates requests by prompt length to reduce Time-to-First-Token (TTFT) latency. It uses a dynamic waiting window and bucketing to group short-prefill requests, improving batch efficiency and CUDA Graph reuse while balancing compute-bound prefill and memory-bound decode phases.
  • FlowPrefill: Addresses head-of-line (HoL) blocking in disaggregated serving by introducing Operator-Level Preemption. This allows fine-grained execution interruption at operator boundaries, decoupling preemption from fixed chunk sizes to balance responsiveness and throughput.
  • Adaptive Rescheduling (ARES) Focuses on the decode phase in disaggregated architectures. It uses an LLM-native predictor to estimate remaining generation length and dynamically migrates requests between decode instances to balance workload variance, reducing P99 token latency by up to 74.77%.
  • General Scheduling Challenges: Research indicates that heterogeneous prefill and decode lengths make minimizing total end-to-end latency NP-hard. Simple heuristics like "shortest-first" fail under variable input sizes, requiring complex quality metrics that balance batch concurrency and response length efficiency under fixed KV-cache memory budgets.
Generated Sep 2, 2026
Open-Weights Reasoning

The paper studies scheduling for long-context LLM serving, where workloads such as retrieval-augmented generation and agentic systems often involve very large prefills followed by short outputs. It focuses on a practical control knob in vLLM: the maximum number of batched tokens (MBT), which determines how many tokens can be admitted into a scheduled batch. The key observation is that MBT does not have a fixed, workload-independent effect on latency; instead, its impact depends on the current scheduling pressure. In other words, the same token-budget setting can be beneficial under some load conditions but harmful under others, especially when prefill-heavy requests compete with short-output requests for scheduling capacity.

The proposed system, P-PAS, introduces prefill-pressure adaptive scheduling to address this issue. Rather than using a static MBT tuned for a particular workload or deployment, P-PAS monitors the pressure of pending prefill work and adapts the token scheduling budget accordingly. The goal is to balance competing objectives: keeping the accelerator sufficiently utilized while avoiding excessive batching that can increase end-to-end latency for latency-sensitive long-context requests. The core contribution is therefore both diagnostic and practical: it identifies MBT as a pressure-sensitive scheduling parameter and proposes an adaptive policy that adjusts it in response to system state.

This matters because long-context serving is becoming a first-class inference pattern, but its latency behavior is less forgiving than shorter prompt workloads. RAG and agentic applications may process tens of thousands of input tokens yet still require fast responses, so throughput-focused scheduling can be insufficient if it inflates request completion time. By making the scheduler aware of prefill pressure, P-PAS offers a way to improve latency predictability and resource efficiency without changing the model itself, and it highlights an underappreciated configuration dimension in modern LLM serving engines.

Generated Sep 2, 2026
Sources