arXiv:2609.14775v1 Announce Type: new Abstract: The rapid rise of generative AI is accelerating cloud data center expansion, with electricity demand projected to double by 2026. Because carbon-intensity varies by more than 5.5x across grids and times of day, where and when inference tasks execute significantly affects operational emissions. We address this issue with three aspects in this paper.
CATS (Carbon-Aware Task Simulator) is a trace-driven discrete-event simulator designed to reduce AI data center emissions by optimizing where and when inference tasks run. It models realistic GPU pools, queuing delays, and SLA constraints to evaluate spatial shifting (routing tasks to cleaner grids) and temporal shifting (delaying tasks for cleaner energy windows).
In a 24-hour simulation of 600,000 tasks across four U.S. grids, the simulator found that spatial shifting reduces CO₂ by 38.4% compared to speed-first scheduling, while temporal shifting achieves a 16.1% reduction with a 3.27% SLA violation rate. The study highlights that 50% of current cloud data centers are located in medium-to-high carbon-intensity grids, indicating significant potential for operational carbon reduction through intelligent scheduling without new infrastructure.
Key limitations noted include that the headline 38.4% savings result is derived from a single 24-hour simulation day, and the workload trace assumes a fixed mix dominated by text generation. The simulator requires real-time marginal carbon-intensity signals (e.g., from WattTime) to function effectively in production environments.
CATS: A Carbon-Aware Task Simulator for Reducing AI Data Center Emissions addresses a growing operational concern in AI infrastructure: the carbon footprint of executing inference workloads in cloud data centers. As generative AI services scale, data center electricity demand is projected to double by 2026, and the paper emphasizes that emissions are not determined solely by energy use. Because grid carbon intensity can vary by more than 5.5× across regions and times of day, the same inference task may produce a materially different amount of CO₂ depending on where and when it runs. CATS is presented as a simulator for studying this problem, likely by modeling workload demand, data center placement, scheduling flexibility, and time-varying grid carbon intensity.
The key contribution is a framework for treating carbon awareness as a first-class dimension in AI workload orchestration. Rather than optimizing only for latency, cost, or utilization, the simulator enables analysis of how task placement and timing affect operational emissions. The central insight is that flexible or deferrable inference workloads can be shifted across data centers or time windows to exploit lower-carbon electricity, potentially reducing emissions without requiring immediate changes to hardware or energy procurement. By quantifying these tradeoffs, CATS can help researchers and operators evaluate carbon-aware scheduling policies before deployment.
This matters because AI data centers are becoming a significant and rapidly growing source of electricity demand, and their environmental impact will depend heavily on operational decisions made at the workload level. A simulator like CATS provides a controlled way to compare scheduling strategies, assess the impact of grid variability, and identify opportunities where carbon reduction can be achieved through software and orchestration rather than solely through hardware efficiency or renewable energy sourcing. In a context of expanding AI infrastructure, such tools are important for aligning performance objectives with lower-emissions operations.