arXiv:2510.05111v2 Announce Type: replace Abstract: We present Agora, a lightweight system designed for cloud-based GPUs which meters a three-dimensional resource vector from commodity performance counters and prices it with a function that charges customers based on resources used rather than time spent renting. Because the calibration is revenue-neutral, any gain must come from demand the flat
Agora is a lightweight system for cloud-based GPUs that implements consumption-based pricing (also called feature-based pricing) by metering a three-dimensional resource vector from commodity performance counters. Unlike traditional flat-rate or time-based billing, Agora charges customers based on actual resource utilization rather than wall-clock time, addressing the economic disconnect where workloads with similar compute needs may consume vastly different amounts of memory bandwidth and interconnect traffic.
The system measures three specific dimensions: tensor throughput (arithmetic rate), HBM bandwidth (device memory data rate), and NVLink traffic (scale-up interconnect rate). By pricing these separately, Agora ensures that bandwidth-bound workloads (such as Large Language Model inference) are not subsidized by compute-bound ones, creating a more equitable and efficient market. The pricing function is calibrated to be revenue-neutral against existing flat-rate models while remaining monotonic to prevent customers from lowering bills by consuming more resources.
Agora’s architecture samples GPU metrics at high frequencies (e.g., ~10 Hz via DCGM) and stores encrypted, itemized logs that allow tenants to verify their bills against their own workload observations. Evaluations on 8x B200 nodes show that workloads can differ by up to 45x on individual resource dimensions, making static tiering insufficient. The system demonstrates that consumption-based pricing can increase provider revenue in 66% of tested configurations and is particularly beneficial when serving 15–30% of the addressable market.
Agora proposes a lightweight consumption-based pricing system for cloud GPUs that shifts billing away from flat, time-based rental toward charges based on the resources a workload actually consumes. The system meters a three-dimensional resource vector using commodity performance counters already available on standard GPUs, avoiding the need for proprietary telemetry or intrusive instrumentation. It then applies a pricing function to that vector, producing a cost that reflects measured usage rather than wall-clock occupancy. This design is aimed at a common mismatch in GPU cloud markets: customers pay for renting a device, but the economic value and operational cost of the workload depend heavily on how intensively its compute, memory, and other resources are used.
A key contribution is the revenue-neutral calibration of the pricing model. By calibrating the consumption function so that provider revenue is preserved under existing demand patterns, Agora frames the potential benefit as coming from demand-side improvements rather than from simply raising effective prices. In other words, if customers or workloads respond to more accurate price signals—by selecting better-matched instances, reducing idle or inefficient usage, or shifting to cheaper resource profiles—the provider can gain from improved utilization and more efficient allocation. This makes the approach attractive as a lower-risk alternative to flat hourly pricing, because it can be deployed without immediately disrupting revenue while still exposing finer-grained economic incentives.
The work matters because GPU workloads are highly heterogeneous: inference, training, fine-tuning, and batch jobs can differ dramatically in utilization and resource intensity, yet many cloud offerings price them with the same coarse time-based model. A consumption-based system can make costs more predictable for variable workloads, reduce the subsidy effect of low-utilization jobs, and encourage resource-aware optimization. For cloud operators, it also offers a path toward more granular monetization of GPU capacity, potentially improving utilization and enabling new service tiers without requiring major hardware or platform changes.