arXiv:2603.28790v2 Announce Type: replace Abstract: In edge computing, the stochastic and bursty nature of serverless workloads challenges autonomous resource orchestration. Traditional reactive controllers, such as the Kubernetes Horizontal Pod Autoscaler (HPA), suffer from reaction latency, leading to Service Level Objective (SLO) violations during traffic spikes and resource flapping during ra
Learning to Remember: Attentive Reinforcement Learning for Edge Serverless Autoscaling (arXiv:2603.28790v2) introduces a stability-aware autoscaling framework that mitigates the "temporal blindness" of standard Deep Reinforcement Learning agents in non-Markovian edge environments. The authors propose an Attention-Enhanced Double-Stacked LSTM architecture integrated within a Proximal Policy Optimization (PPO) agent to unify short-term temporal context with control decisions.
Key technical and performance highlights include:
The computational overhead of the attention-based network is negligible, measuring only 0.70 ms on CPU and 0.80 ms on GPU per 60-second control step, representing approximately 0.001% of the control interval.