arXiv:2609.20874v1 Announce Type: cross Abstract: Large language model (LLM) inference deployed on Kubernetes faces an autoscaling challenge that conventional web services do not: new serving replicas take two to ten minutes to start because multi-gigabyte model weights must be loaded, which makes purely reactive scaling structurally late. We ask a sharp question: among the components of a predic
Decomposing Predictive Kubernetes Autoscaling for Large Language Model Serving Under Long Startup Delays (arXiv:2609.20874v1) is a paper accepted by CloudCom 2026 that addresses the challenge of autoscaling LLM inference on Kubernetes, where replica startup delays of two to ten minutes make reactive scaling ineffective.
The study decomposes predictive autoscaling into four factors to identify which components matter most under these long actuation delays:
Key findings indicate that a simple Exponentially Weighted Moving Average (EWMA) predictor with delay-aware lookahead and a UCB margin is superior to more complex models like Kalman filters, reducing Time-to-First-Token (TTFT) violations from 53% (reactive) to 0.5%. Real-cluster validation on a Qwen2.5-7B model confirmed that predictive lookahead cuts TTFT violations from 63.5% to 3.7% compared to reactive KEDA scaling.
The paper examines autoscaling for LLM inference on Kubernetes, where the dominant failure mode of conventional reactive scaling is not just forecast error but the long startup time of new replicas. Because multi-gigabyte model weights must be loaded before a serving instance can accept traffic, a new pod may take two to ten minutes to become useful. In that regime, a controller that reacts only after queueing or latency violations are already visible is structurally late: by the time it decides to scale out, the workload may have already degraded, and the added capacity arrives too slowly to prevent the violation.
Its central contribution is to decompose predictive Kubernetes autoscaling into its constituent parts rather than treating it as a single forecasting problem. The work focuses on which components matter—demand prediction, startup-time estimation, lead-time-aware capacity planning, and safe actuation under uncertainty—and how they interact when scaling actions are slow and expensive. The key insight is that for LLM serving, prediction quality alone is insufficient; the autoscaler must explicitly reason about provisioning delay, GPU resource cost, batching behavior, and the asymmetric consequences of over- versus under-provisioning.
This matters because LLM workloads are typically GPU-intensive, bursty, and latency-sensitive, making them a poor fit for web-service autoscaling assumptions. Understanding the decomposition of predictive scaling helps platform teams design more realistic control policies, such as warm capacity pools, model-weight caching, preloading, or lead-time-aware scaling rules. More broadly, the paper highlights that cloud-native serving systems need autoscalers that are not only predictive, but also delay-aware and cost-aware.