arXiv:2603.28790v2 Announce Type: replace Abstract: In edge computing, the stochastic and bursty nature of serverless workloads challenges autonomous resource orchestration. Traditional reactive controllers, such as the Kubernetes Horizontal Pod Autoscaler (HPA), suffer from reaction latency, leading to Service Level Objective (SLO) violations during traffic spikes and resource flapping during ra

Topological visualization of Learning to Remember: Attentive Reinforcement Learning for Edge Serverless Autoscaling
Brave API

Learning to Remember: Attentive Reinforcement Learning for Edge Serverless Autoscaling (arXiv:2603.28790v2) introduces a stability-aware autoscaling framework that mitigates the "temporal blindness" of standard Deep Reinforcement Learning agents in non-Markovian edge environments. The authors propose an Attention-Enhanced Double-Stacked LSTM architecture integrated within a Proximal Policy Optimization (PPO) agent to unify short-term temporal context with control decisions.

Key technical and performance highlights include:

  • Architecture: The agent constructs a 13-dimensional state vector using the 8 most recent observation vectors as input. A learned soft-attention mechanism weights these historical states non-uniformly, suppressing high-frequency jitter while preserving trends preceding demand shifts.
  • Kubernetes-Native Control: Unlike models that output raw replica counts, this supervisory controller maps actions to HPA CPU utilization targets (discrete set: 30%, 50%, 70%, 90%) and gateway rate limits, preserving Kubernetes' fail-safe mechanisms.
  • Latency Reduction: The framework cuts P90 tail latency by 67% compared to a single-layer LSTM ablation and holds average latency within a 50ms hard SLO for 98.8% of the run, compared to 49.6% for the ablation and 43.5% for static HPA.
  • Stability and Efficiency: It matches the latency performance of Kubernetes Event-Driven Autoscaling (KEDA) but with 75% fewer replica-steps and 59% less replica churn. The agent prevents premature scale-downs during transient dips, reducing SLO violation bursts to a maximum of 5 consecutive intervals (vs. 24 for KEDA).
  • Validation: The study was conducted on two independent heterogeneous Kubernetes clusters using real-world Azure Functions traces, with statistical significance confirmed via matched-pair testing (Wilcoxon, p=0.0020).

The computational overhead of the attention-based network is negligible, measuring only 0.70 ms on CPU and 0.80 ms on GPU per 60-second control step, representing approximately 0.001% of the control interval.

Generated 10d ago
Open-Weights Reasoning
  • Problem focus: The paper targets autonomous autoscaling for serverless workloads in edge computing environments, where traffic is highly stochastic and bursty. It argues that conventional reactive controllers—most notably the Kubernetes Horizontal Pod Autoscaler (HPA)—are poorly matched to these dynamics because they rely on observed utilization or request metrics and therefore incur reaction latency. That lag can cause SLO violations during sudden traffic spikes and induce resource flapping when workloads change rapidly.
  • Core contribution: The work proposes an attentive reinforcement learning approach to edge serverless autoscaling, framing capacity provisioning as a sequential decision-making problem under uncertainty. The “learning to remember” aspect suggests that the controller uses attention or memory mechanisms to weigh relevant historical observations, enabling it to learn temporal workload patterns and make more anticipatory scaling decisions rather than merely reacting to current load signals. The key insight is that edge serverless autoscaling benefits from predictive, context-aware policy learning rather than fixed threshold-based feedback control.
  • Why it matters: This is relevant because edge platforms often operate under tighter latency, cost, and resource constraints than centralized clouds, and serverless workloads can be highly irregular. An RL-based autoscaler that can anticipate bursts and reduce oscillatory scaling behavior could improve SLO attainment, resource efficiency, and operational stability. The paper is therefore significant for researchers and practitioners interested in adaptive resource orchestration for edge-native serverless systems.
Generated 10d ago
Sources