arXiv:2609.29160v1 Announce Type: new Abstract: Multi-model LLM serving is moving toward shared MaaS clusters, where co-hosted models compete for a fixed GPU budget while each model experiences time-varying demand and must satisfy its own latency SLO. Existing LLM autoscalers remain largely model-local: their signals expose local runtime activity or delayed latency outcomes, and their scale-up or
Cross-Model Autoscaling for Shared LLM Serving (arXiv:2609.29160v1, submitted September 24, 2026) introduces TRE (Token-service-share Rebalancing Engine), a control-plane framework designed to optimize shared Model-as-a-Service (MaaS) clusters with fixed GPU budgets.
The material addresses multi-model LLM serving in shared model-as-a-service (MaaS) clusters, where multiple language models are co-hosted on a fixed GPU pool while facing independent, time-varying demand and per-model latency SLOs. It frames autoscaling not as a collection of isolated, model-local control loops, but as a cluster-level coordination problem. In such environments, existing autoscalers often rely on local runtime signals or delayed latency feedback, which can be noisy, slow to react, and blind to contention between co-located models.
Its key insight is that effective capacity management in shared LLM serving requires cross-model visibility: decisions for one model should account for pressure on other models and on the shared GPU budget. The work therefore proposes a cross-model autoscaling approach that coordinates scale-up and scale-down behavior across models rather than treating each workload in isolation. The goal is to make scaling more proactive and globally aware, reducing SLO violations under bursty or uneven demand while improving overall GPU utilization.
This matters because modern LLM platforms are increasingly consolidating many models onto common serving infrastructure, making per-model scaling insufficient on its own. The paper is relevant to LLM serving systems, MaaS operators, and SRE teams that need to balance latency guarantees, resource efficiency, and multi-tenant fairness in a competitive GPU environment.