Concurrent multi-agent workflows expose future dependencies and serving-state requirements while running on heterogeneous GPU pools with time-varying load, model residency, and resource availability. The logical workflow defines the required computation, whereas its physical scheduling units, model-lifecycle actions, resource ordering, and placement must be selected according to the observed pool
Latency-Aware Orchestration for Multi-Agent LLM Workflows on Heterogeneous GPUs is a prediction-guided runtime system introduced in September 2026 by Jinghao Wang and colleagues to optimize concurrent multi-agent LLM execution. The system addresses the challenge of managing future dependencies and serving-state requirements in heterogeneous GPU pools by separating logical workflow semantics from state-dependent physical execution.
It employs three core components: a Predictor to estimate device-specific activation latency, memory, and loading costs; a Constructor to build semantics-preserving fusion and model-lifecycle alternatives; and a Scheduler to jointly optimize selection, placement, and execution order based on live pool state. This approach allows the system to prefetch deployments, fuse same-deployment activations, and coordinate compatible model lifecycles across workflows.
Experimental results demonstrate that this orchestration method significantly improves performance over state-of-the-art schedulers like Parrot and Kairos. Specifically, it reduces burst makespan by up to 36.8% and overall p95 completion latency by 25.9%, while saving up to 24.63 GPU-s per completed session across varied workflow scenarios.
Latency-Aware Orchestration for Multi-Agent LLM Workflows on Heterogeneous GPUs examines how to execute concurrent multi-agent LLM workloads efficiently on GPU pools that are heterogeneous, dynamically loaded, and stateful. Rather than treating a workflow as a simple sequence of inference calls, it distinguishes the logical workflow—what computation the agents require—from the physical execution decisions that must be made at runtime: which scheduling units to issue, when to load or evict models, how to order resource requests, and where to place work given current residency, capacity, and contention.
A central contribution is a latency-aware orchestration perspective that jointly considers workflow dependencies, model-lifecycle actions, and pool state. By exposing future dependencies and serving-state requirements while the workflow is running, the approach can make scheduling and placement decisions that reduce unnecessary model migrations, avoid resource bottlenecks, and improve end-to-end response time. This is important because multi-agent LLM systems often involve many interleaved requests, repeated model switches, and uneven GPU demand, where naive per-request or per-agent scheduling can lead to poor tail latency and underutilized accelerators.
Overall, the material matters because it targets a practical gap in serving complex LLM applications: production multi-agent workloads are not just a collection of independent prompts, but structured, stateful, and resource-sensitive processes. By modeling orchestration as a coordinated problem over workflow semantics, model lifecycle, and heterogeneous GPU availability, the work provides a more realistic foundation for building low-latency, cost-efficient LLM serving systems in shared or dynamic GPU environments.
This arXiv paper addresses the runtime orchestration problem for multi-agent LLM workflows that run on heterogeneous GPU pools. Its central premise is that agent workflows are not merely collections of independent inference requests: even while executing, the logical workflow reveals future dependencies, parallel branches, and serving-state requirements, such as which models must be resident, warm, or co-scheduled to avoid stalls. The paper frames orchestration as a coupled decision problem over physical scheduling units, model-lifecycle actions such as loading, evicting, or migrating models, resource ordering, and placement across GPU devices with differing capabilities and time-varying load.
The key insight is that end-to-end latency in such systems is often dominated by orchestration overhead rather than raw model compute: model residency decisions, hardware fit, and dependency-aware resource ordering can determine whether a workflow makes progress or waits on memory pressure, model switching, or misplacement. By treating the logical workflow and the observed pool state together, the work emphasizes that an effective orchestrator must co-optimize computation scheduling with serving-state management, rather than applying per-request or per-model heuristics in isolation.
This matters because multi-agent LLM systems are increasingly deployed on mixed GPU fleets where memory, interconnect, and utilization are dynamic. A latency-aware orchestration layer can reduce tail latency, improve GPU utilization, and better match workflow structure to hardware heterogeneity. More broadly, the material points toward a systems-level approach to agent execution in which scheduling, model lifecycle, and placement are treated as first-class optimization variables, an important direction as LLM workflows become larger, more interactive, and more resource-constrained.