The rapid growth of LLM inference is shifting sustainability concerns from one-time training to continuous serving, where infrastructure decisions shape energy use, carbon emissions, water consumption, and service quality. Yet operators often need to compare deployment alternatives before large-scale infrastructure is built, making direct measurement costly, slow, and sometimes infeasible. We pres

Topological visualization of InFactPlanner: Planning Sustainable Geo-Distributed LLM Data Centers
Brave API

Recent research highlights that LLM inference now dominates the environmental footprint of large language models, with operational carbon emissions and water usage far exceeding those of the training phase.

Key sustainability challenges include: Water Scarcity: Cloud datacenters consume vast amounts of water, with estimates indicating 500 ml of fresh water is expended for every 20–50 LLM inference requests. Carbon Intensity: Inference can have a carbon footprint 1,400× larger than training annually for large-scale deployments, contributing significantly to global emissions. * Cost Disparity: The cost of operating LLMs in their inference phase can exceed training costs by 25× per year.

Proposed solutions focus on geo-distributed scheduling frameworks: SLIT Framework: A novel approach that co-optimizes Time-to-First-Token (TTFT), carbon emissions, water usage, and energy costs using a machine learning-based metaheuristic. It generates real-time Pareto fronts to allow datacenter managers to balance performance and sustainability. Performance: SLIT significantly outperforms state-of-the-art frameworks like Helix and Splitwise, reducing carbon emissions by up to 99% and water usage by up to 99% while maintaining quality of service. * Comprehensive Modeling: Effective planning requires modeling energy costs, carbon intensity, and water intensity across heterogeneous hardware and global regions (e.g., East Asia, North America, Western Europe, Oceania).

Deployment-aware frameworks also propose using mixed-integer linear programming to enforce Service Level Objectives (SLOs) while minimizing energy and water use. Such policies can reduce per-prompt energy by 57–59%, water by 59–60%, and carbon by 78–80% through batch right-sizing, semantic token control, and phase-aware hardware placement.

Generated Aug 24, 2026
Open-Weights Reasoning

InFactPlanner addresses a growing operational challenge in LLM infrastructure: as model deployment shifts from episodic training to continuous, large-scale inference, sustainability is no longer just a one-time training cost problem but a persistent serving-time decision problem. Geo-distributed data-center placement, hardware selection, traffic routing, and energy sourcing all jointly determine energy consumption, carbon intensity, water use, and service quality. The paper is motivated by the fact that operators often need to compare deployment alternatives before committing to expensive, long-lived infrastructure, where direct measurement is slow, costly, or impossible.

The core contribution is a planning framework for evaluating sustainable geo-distributed LLM data-center designs before they are built. Rather than treating sustainability and performance as separate afterthoughts, the work frames them as coupled planning variables: regional grid conditions, renewable availability, data-center siting, inference workload patterns, latency or SLO constraints, and infrastructure choices all interact. This enables scenario-level comparison of deployment options and helps identify tradeoffs among carbon, energy, water, cost, and user-facing service quality in a way that supports more informed capex and placement decisions.

The paper matters because it targets a decision point where mistakes are expensive and hard to reverse. As inference workloads become a dominant share of LLM operational load, the ability to model and compare geo-distributed architectures in advance is essential for operators seeking to reduce environmental impact without sacrificing reliability or performance. In that sense, the work is less about a single optimization result and more about providing a practical, pre-deployment lens for designing LLM serving infrastructure that is both operationally competitive and sustainability-aware.

Generated Aug 24, 2026
Sources