Achieving cost efficiency while meeting strict user-facing SLOs (e.g., time-to-first-token) remains a fundamental challenge for cloud GPU clusters serving large language models (LLMs). Autoscaling is the key mechanism for cluster resource management, yet a basic system design question is open for serving LLMs: what should be the unit of scaling? Existing approaches primarily treat the entire model
OpScale (Operator-level Autoscaling) is a framework that shifts resource management granularity from the entire model to individual operators (e.g., attention, linear transformation) within large generative models.
OpScale tackles a core systems question in LLM serving: what is the right granularity for autoscaling? As cloud GPU clusters increasingly serve large language models under strict user-facing SLOs such as time-to-first-token, coarse-grained scaling—typically adding or removing whole model replicas or serving instances—can be inefficient. Different parts of the inference pipeline experience different load dynamics: prefill, decode, attention, feed-forward, and other operator-level stages may become bottlenecks at different times depending on request mix, sequence length, batching behavior, and model architecture. The paper argues that treating the entire model as the unit of scaling leaves performance and cost tradeoffs on the table.
Its key contribution is an operator-level provisioning and autoscaling framework that decomposes LLM serving into finer-grained components and scales them more independently. Rather than only asking “how many copies of the model should we run?”, OpScale asks “which operators or stages need more GPU capacity right now?” and adjusts resources accordingly. This enables the system to respond more precisely to the actual bottleneck in the serving pipeline, allocate GPU capacity where it is needed, and avoid over-provisioning stages that are not currently limiting SLO attainment.
The work matters because it reframes autoscaling for LLM serving from a monolithic, instance-centric problem into a more nuanced resource-management problem. If operator-level scaling can improve GPU utilization while preserving latency SLOs, it offers a practical path toward lower serving costs at scale. It also aligns with broader trends in LLM infrastructure, including disaggregated prefill/decode serving, heterogeneous accelerators, and increasingly complex model architectures where different components have different compute, memory, and latency characteristics. In short, OpScale provides a systems-level insight: the unit of scaling should match the granularity of the workload’s bottleneck, not just the boundary of the deployed model.