arXiv:2609.26763v1 Announce Type: cross Abstract: Recent advances in large language models (LLMs) are driving the emergence of multi-modal and agentic services for mobile users through cloud and edge infrastructures, where long-context workloads pose daunting challenges for inference latency. Existing disaggregated LLM serving systems largely rely on hardware profiling, configuration enumeration,

Topological visualization of SARA: SLO-Aware Resource Allocation for Disaggregated Agentic LLM Services
Brave API

SARA is a resource allocation framework for disaggregated agentic LLM services that maximizes system goodput under deployment cost and Service Level Objective (SLO) constraints. By modeling prefill, KV cache transfer, and decode stages using queuing theory, it derives analytical resource requirements based on workload characteristics and hardware parameters.

Key performance metrics include: Accuracy: Predicts stage-wise SLOs with mean errors below 5%. Efficiency: Improves system goodput by 26.6% on average compared to state-of-the-art baselines under identical cost constraints. * Scalability: Handles both light- and heavy-tailed workloads, revealing that prefill and decode device requirements scale quadratically with input and output lengths.

The framework was submitted to arXiv in September 2026 by Shicong Liu, Xianghao Yu, Zhen Gao, and Jun Zhang.

Generated 11d ago
Open-Weights Reasoning

The paper addresses resource management for disaggregated LLM serving in the context of agentic, multi-modal services deployed across cloud and edge infrastructures. Such workloads are especially demanding because agentic LLM applications often involve long contexts, repeated inference calls, tool use, and end-to-end latency objectives rather than a single one-shot completion. In disaggregated serving, stages such as prefill and decode can be placed on different accelerators or nodes, which improves flexibility but makes resource allocation harder: the system must decide how much compute, memory, and network capacity to assign to each stage and each request class while still meeting service-level objectives. The paper positions itself against existing disaggregated serving systems that often depend on hardware profiling and configuration enumeration, approaches that can be costly, brittle, and difficult to scale as models, context lengths, and agent workloads evolve.

Its central contribution is SARA, an SLO-aware resource allocation framework for disaggregated agentic LLM services. Rather than treating serving configuration as a static, manually tuned problem, SARA frames allocation as an optimization problem conditioned on the desired latency or quality targets. It is designed to handle the variable prefill/decode demands characteristic of long-context agentic workloads and to make resource decisions that balance SLO satisfaction against efficiency metrics such as throughput, utilization, or infrastructure cost. The key insight is that agentic mobile and edge services need a serving layer that is not only high-performance in isolation, but also predictable under heterogeneous workloads and resource constraints.

This matters because agentic LLM services are moving toward latency-sensitive, multi-step user experiences where poor tail latency or overprovisioning can undermine both usability and cost efficiency. By focusing on SLO-aware allocation in disaggregated settings, the work is relevant to operators trying to serve large, context-heavy models on shared cloud/edge infrastructure without relying on exhaustive offline tuning. More broadly, it points toward a needed shift in LLM serving systems: from static, hardware-centric configuration toward adaptive, workload-aware resource management that can keep pace with increasingly complex agentic applications.

Generated 11d ago
Sources