arXiv:2609.39067v1 Announce Type: cross Abstract: Elastic Compute Cloud (EC2) Spot is 60% to 90% cheaper than On-Demand but can be reclaimed on just a 2-minute notice; for expensive multi-node training this loss can be severe, with one reclaim costing hours of synchronous progress. We build Argus, a Kubernetes operator, and ask empirically, on a CIFAR-10 testbed, when predicting interruptions bea

Topological visualization of Argus: A Real-EKS Study of When Predicting Spot Interruptions Beats Simple Checkpointing
Brave API

Argus is a Kubernetes operator designed to optimize AWS EC2 Spot instance usage for multi-node training by predicting interruptions to trigger proactive checkpointing. While Spot instances offer 60% to 90% cost savings, they can be reclaimed with only a 2-minute warning, risking significant progress loss in expensive jobs.

The study found that predictive checkpointing achieves zero wasted compute by migrating nodes before interruption, whereas reactive (2-minute notice) or simple periodic checkpointing still incur waste or inefficiency. However, predictive methods require careful lead-time tuning; a small lead time suffices for zero waste, but excessive lead causes over-migration and reduced completion rates.

Key findings from the 80-trial benchmark on a CIFAR-10 testbed include: Predictive Policy: 0.0 s wasted compute, but higher makespan (353.2 s) and checkpoint count (221.6). Periodic Policy: Low waste (4.3 s), significantly faster makespan (158.8 s), and fewer checkpoints (24.0). * Reactive Policy: Higher waste (169.4 s) as the 2-minute notice becomes ineffective when interruptions occur faster than once every two minutes.

The system uses a Transformer risk model trained on proxy labels (Spot price spikes) to advise checkpointing, demonstrating that prediction beats simple checkpointing only when the prediction lead aligns closely with the interruption frequency.

Generated 3d ago
Open-Weights Reasoning

Argus addresses a practical but consequential problem in cost-effective cloud training: EC2 Spot instances can be 60–90% cheaper than On-Demand instances, but they can be reclaimed with only a two-minute notice. For multi-node synchronous training, a single interruption can discard substantial progress, making the choice of interruption-mitigation strategy important. The paper studies this problem on real Amazon EKS and builds Argus, a Kubernetes operator designed to manage training workloads under Spot reclamation risk. Its central empirical question is whether predicting Spot interruptions provides a meaningful advantage over a simpler strategy such as periodic or “simple” checkpointing.

The key contribution is an empirical comparison, using a CIFAR-10 testbed, of predictive interruption handling versus straightforward checkpointing. Rather than assuming that better prediction is always beneficial, the study examines when prediction actually pays off. That framing matters because prediction systems can introduce false positives, extra checkpointing overhead, scheduling complexity, and wasted work if the model is conservative. By contrast, simple checkpointing is easier to operate but may checkpoint too frequently or fail to preserve enough state when interruptions arrive. The paper’s insight is that the value of prediction depends on the cost of checkpointing, the predictability of interruptions, the granularity of training progress, and the overhead of acting on a prediction.

This work matters because it moves the discussion from “use Spot to save money” to a more nuanced systems question: how should a cluster operator trade off prediction accuracy, checkpoint cost, and training throughput? For organizations running expensive multi-node jobs on Kubernetes, the result is a practical perspective on when a prediction-driven operator like Argus can outperform a simpler checkpointing baseline, and when the added complexity may not be justified. In short, it provides evidence-based guidance for designing cost-aware training infrastructure on Spot fleets.

Generated 3d ago
Sources