LLM-agent services repeatedly execute small deterministic transitions between model and tool calls: route an outcome, update state, and emit the next effect. We ask when this control path exposes enough concurrent work for GPU execution, and what changes when a GPU-computed route decision remains on device. We formalize the ready-cohort boundary using fixed-partition share F, exact offline share P
Ready Cohorts is a framework that identifies two measurable gates for executing LLM-agent control paths on GPUs: deadline-feasible cohort supply and observation placement.
The study concludes that a GPU control plane is only viable if it can pass both gates, requiring an online route compactor that batches events and maintains device-resident decisions to avoid CPU-GPU synchronization overhead.
The paper examines a performance problem that is easy to overlook in LLM-agent systems: the control path between model calls and tool calls. Many agent workloads are not dominated only by large inference calls; they also contain frequent, small, deterministic transitions—routing an outcome, updating execution state, and emitting the next effect. These steps are often short, dependency-heavy, and latency-sensitive, so the question is not whether they are “compute-heavy” in isolation, but whether they can be organized into enough concurrent work to make GPU execution worthwhile. The paper frames this as a scheduling and offload question: when does the agent control loop expose a sufficiently large batch of ready work, and how does that change if the routing decision itself is computed on the GPU instead of being returned to the host?
Its central contribution is a formal notion of a ready cohort: a set of agent-control transitions that are simultaneously executable because their dependencies have been satisfied and their next actions are well-defined. The authors use two parameters to bound the GPU opportunity: a fixed-partition share F, which captures work that can be assigned under a static or predictable partitioning, and an exact offline share P, which captures work whose readiness can be determined more precisely from offline analysis. Together, these quantities let the paper reason about the fraction of the agent loop that can be treated as GPU-ready work, rather than assuming that all non-inference logic is either trivial or offloadable. A key insight is that the boundary is not just about total compute, but about concurrency, dependency locality, and whether control decisions can be resolved without crossing the CPU–GPU boundary.
The material matters because LLM-agent services are increasingly structured as tight interleavings of model inference, tool execution, and stateful control logic. In such systems, host round trips can become a hidden bottleneck: even if each routing or state-update step is small, repeatedly returning control to the CPU can add latency, reduce parallelism, and prevent the GPU from keeping its execution units busy. By showing how to bound the ready-cohort opportunity and by contrasting host-mediated control with on-device GPU routing, the paper gives system designers a more precise way to decide when GPU offload of agent control is beneficial. More broadly, it reframes agent-system performance as a problem of control-flow batching and device-resident decision making, not merely as a problem of faster model inference.