arXiv:2609.19936v1 Announce Type: new Abstract: Memory over-provisioning results in resource underutilization when HPC workloads run on Kubernetes. The default Vertical Pod Autoscaler (VPA) cannot anticipate phase-driven memory spikes for first-run HPC jobs. In this work, we present a reinforcement learning (RL) recommender VERA that formulates vertical memory scaling as a Markov Decision Process

Topological visualization of VERA: Reinforcement Learning for Dynamic Memory Scaling of HPC Workloads in Kubernetes
Brave API

VERA is a reinforcement learning (RL) recommender designed to dynamically scale memory for High-Performance Computing (HPC) workloads on Kubernetes, addressing the default Vertical Pod Autoscaler's (VPA) inability to anticipate phase-driven memory spikes. Formulated as a Markov Decision Process, VERA utilizes a PPO agent trained on 3,353 real Prometheus traces to autonomously adjust memory limits using a continuous action space and a phase-aware reward structure.

Evaluated on a live Google Kubernetes Engine cluster across workloads like LAMMPS and MLPerf, VERA reclaims 31.6% of available memory headroom while incurring at most one out-of-memory (OOM) event per 100 runs. In contrast, the default VPA increased memory provisioning by 7.9% and its recommendations would have been insufficient to avoid OOMs in 30 runs, demonstrating VERA's superior ability to reduce waste and prevent crashes in dynamic HPC environments.

Generated 14d ago
Open-Weights Reasoning

Problem and approach. The paper addresses memory over-provisioning for HPC workloads running on Kubernetes, where applications often exhibit sharp, phase-driven memory spikes that are hard to predict from static resource requests. The default Vertical Pod Autoscaler (VPA) is reactive and relies on observed utilization, which makes it poorly suited to first-run HPC jobs that lack prior execution history. VERA instead formulates vertical memory scaling as a Markov Decision Process, using reinforcement learning to recommend memory scaling decisions as the workload evolves.

Contribution and significance. The key contribution is an RL-based recommender that treats memory allocation as a sequential decision problem rather than a one-shot sizing problem. This allows VERA to reason about temporal workload behavior and phase transitions that are common in HPC applications, rather than relying solely on historical baselines or post-hoc utilization feedback. This matters because it targets a core weakness of containerized HPC deployment: balancing memory efficiency against the risk of out-of-memory failures. If it improves first-run and phase-driven scaling, VERA could reduce idle memory, increase node packing efficiency, and make Kubernetes a more practical platform for memory-intensive HPC workloads.

Generated 14d ago
Sources