Load shifting across geographic regions to chase intermittent renewable energy availability is commonly used in reducing cloud infrastructure carbon footprint. However, it often omits low-latency applications due to high latency variances of wide area networks (WAN) that interconnect regions. This paper addresses accommodating low-latency applications into load shifting by minimizing their shiftin

Topological visualization of A Technique for Load Shifting Low-latency Applications in Multi-Region Renewables Harvesting via SMT Core Pooling
Brave API

A Technique for Load Shifting Low-latency Applications in Multi-Region Renewables Harvesting via SMT Core Pooling is a September 2026 paper by Tharindu B. Hewage, Shashikant Ilager, Maria A. Rodriguez, and Rajkumar Buyya (arXiv:2609.03297).

The authors address the limitation that load shifting across geographic regions to chase renewable energy often excludes low-latency applications due to high latency variances from wide area networks (WAN). To resolve this, they propose a hardware-software co-design that uses Simultaneous Multi-Threading (SMT) core pooling to maintain a static set of logical cores within local cloud regions during renewable energy supply valleys.

Key outcomes and features include: Hardware Mechanism: By deep-idling physical cores and leveraging SMT, the system ensures static CPU capacity for low-latency workloads, eliminating the need to offload them across WANs. Software Scheduling: A novel VM scheduling algorithm prioritizes low-latency VMs for local placement while offloading best-effort applications to other regions when necessary. * Performance Results: The technique achieved an 80% reduction in offloading low-latency VMs, a 43.81% reduction in the coefficient of variation of p90 end-user latency, and a worst-case latency compromise of only 11.97% due to SMT core usage.

Generated 28d ago
Cerebras Thinking

This material addresses a key limitation in carbon-aware cloud scheduling: while shifting workloads across geographic regions to follow intermittent renewable energy can reduce datacenter carbon emissions, low-latency applications have generally been excluded because wide-area network conditions introduce high and variable tail latency. The paper focuses on closing that gap by asking how latency-sensitive workloads can be included in multi-region renewable-energy harvesting without violating strict latency service-level objectives.

Its central technique is SMT core pooling, which uses simultaneous multithreading hardware to create a more flexible yet controllable compute resource pool. Rather than treating SMT logical cores as fully interchangeable, the approach manages how sibling threads are shared, reserved, or pooled across regions so that shifted low-latency workloads can obtain predictable compute capacity and bounded interference. In effect, the system couples region-level energy and network-aware load shifting with core-level scheduling control to minimize the latency penalty introduced by moving latency-sensitive tasks.

The work matters because it extends carbon-aware load shifting beyond batch and elastic workloads into a much larger class of production workloads that are currently locked to fixed regions for latency reasons. By showing how SMT-level resource management can make low-latency shifting safer, the paper offers a practical path toward broader renewable-energy utilization in multi-region cloud infrastructure while preserving user-facing performance guarantees.

Generated 28d ago
Open-Weights Reasoning

The paper addresses a key limitation of renewable-energy-aware cloud scheduling: load shifting workloads across geographic regions to follow intermittent solar and wind generation is a common carbon-reduction strategy, but it often excludes low-latency applications because wide-area network paths between regions can introduce large and unpredictable latency variance. The authors propose a technique that makes latency-sensitive workloads more amenable to load shifting by explicitly minimizing the latency disruption caused by moving them. The central mechanism is SMT core pooling, in which simultaneous multithreading cores or hardware threads are pooled and dynamically allocated to create latency-aware capacity headroom in candidate regions. This allows the system to place or migrate low-latency applications only when the target region can absorb them within an acceptable latency budget, while packing or deferring less latency-critical work onto the pooled SMT capacity.

The main insight is that including low-latency applications in renewable load shifting is not only a network problem; it is also a local resource-isolation and scheduling problem. By treating SMT cores as a flexible pool rather than fixed, homogeneous CPU capacity, the technique can provide rapid, fine-grained capacity for latency-critical tenants and reduce tail-latency risk during cross-region shifts. The contribution is therefore a hardware-aware placement/scheduling mechanism that couples renewable availability, regional compute capacity, and latency SLOs, rather than simply maximizing renewable energy use or shifting all workloads uniformly.

This matters because latency-sensitive workloads can represent a substantial share of cloud traffic, and excluding them limits the carbon savings achievable through multi-region renewable harvesting. If operators can shift low-latency applications safely, they can exploit more renewable energy across a broader workload mix without violating user-facing performance commitments. The work is therefore relevant to sustainable cloud design, multi-region resource orchestration, and scheduling systems that must balance energy, cost, and tail-latency objectives.

Generated 28d ago
Sources