arXiv:2609.29160v1 Announce Type: new Abstract: Multi-model LLM serving is moving toward shared MaaS clusters, where co-hosted models compete for a fixed GPU budget while each model experiences time-varying demand and must satisfy its own latency SLO. Existing LLM autoscalers remain largely model-local: their signals expose local runtime activity or delayed latency outcomes, and their scale-up or

Topological visualization of Cross-Model Autoscaling for Shared LLM Serving
Brave API

Cross-Model Autoscaling for Shared LLM Serving (arXiv:2609.29160v1, submitted September 24, 2026) introduces TRE (Token-service-share Rebalancing Engine), a control-plane framework designed to optimize shared Model-as-a-Service (MaaS) clusters with fixed GPU budgets.

  • Core Mechanism: TRE utilizes Token Service Share (TSS), a demand-normalized signal that enables cross-model comparison of effective token service, allowing the system to rank service deficits across heterogeneous models and SLO classes.
  • Operation: The framework coordinates bounded receiver–donor capacity movement under a fixed budget, separating fast rescue operations from slower rebalancing to incrementally reallocate active replicas toward models with the largest calibrated service deficits.
  • Performance: Evaluated on a Kubernetes-based AIBrix and vLLM stack, TRE reduces P95 end-to-end latency by 11.9–79.0% and P99 latency by 12.5–72.6% compared to state-of-the-art KV-cache-based reactive autoscalers across seven LLM serving traces.
  • Impact: This approach addresses the limitation of existing model-local autoscalers, which fail to arbitrate shared capacity contention when the cluster is fully occupied, thereby maximizing overall SLO attainment in multi-model environments.
Generated 7d ago
Open-Weights Reasoning

The material addresses multi-model LLM serving in shared model-as-a-service (MaaS) clusters, where multiple language models are co-hosted on a fixed GPU pool while facing independent, time-varying demand and per-model latency SLOs. It frames autoscaling not as a collection of isolated, model-local control loops, but as a cluster-level coordination problem. In such environments, existing autoscalers often rely on local runtime signals or delayed latency feedback, which can be noisy, slow to react, and blind to contention between co-located models.

Its key insight is that effective capacity management in shared LLM serving requires cross-model visibility: decisions for one model should account for pressure on other models and on the shared GPU budget. The work therefore proposes a cross-model autoscaling approach that coordinates scale-up and scale-down behavior across models rather than treating each workload in isolation. The goal is to make scaling more proactive and globally aware, reducing SLO violations under bursty or uneven demand while improving overall GPU utilization.

This matters because modern LLM platforms are increasingly consolidating many models onto common serving infrastructure, making per-model scaling insufficient on its own. The paper is relevant to LLM serving systems, MaaS operators, and SRE teams that need to balance latency guarantees, resource efficiency, and multi-tenant fairness in a competitive GPU environment.

Generated 7d ago
Sources