Systematic review of 66 LLM studies for HVAC operations (2023–2026), classifying them into application and method families while assessing realism and deployment readiness.
Recent systematic reviews indicate that Large Language Models (LLMs) are transitioning from experimental tools to core infrastructure in building energy applications, with significant focus on HVAC operations, fault detection, and human-centric thermal comfort. While LLMs excel at semantic data integration, automated fault diagnosis, and natural language-based occupant feedback interpretation, they currently lack the stability and precision required for high-frequency, low-level control execution, which remains better suited for conventional Model Predictive Control (MPC) or Reinforcement Learning (RL) methods.
Deployment readiness is currently limited by challenges in real-time reliability, safety verification, and the lack of standardized benchmarks for field testing. Most existing studies are simulations or short-term proof-of-concepts; real-world deployments typically show single- to low-double-digit energy savings, and critical gaps remain in long-term validation, explainability under uncertainty, and edge-deployment constraints. Future advancements require hybrid architectures that combine LLMs for high-level reasoning and semantic orchestration with physics-grounded, validated controllers for deterministic thermal management.
Key findings from recent literature (2023–2026) include:
This material is a systematic review of 66 studies from 2023–2026 that apply large language models to HVAC operations in building energy systems. It covers a broad set of use cases, including control and scheduling, energy management, fault detection and diagnosis, occupancy-aware optimization, and operator-facing interfaces. The review organizes the literature into application families and method families, distinguishing approaches such as prompt-based reasoning, retrieval-augmented generation over building data or equipment documentation, fine-tuned domain models, and agentic or tool-augmented pipelines that may interface with building automation systems, weather data, or optimization solvers. A central contribution is its deployment-readiness framing: it evaluates not only whether an LLM can generate plausible text, but whether the study reflects realistic building conditions, physically valid actions, safety constraints, validation against conventional baselines, and operational concerns such as latency, auditability, human oversight, and integration with existing BMS/BAS stacks.
The key insight is that LLMs are most promising as interpretable reasoning and orchestration layers, rather than as drop-in replacements for physics-based controllers or established energy optimization methods. The strongest reviewed work tends to combine language-model reasoning with grounded domain data, constrained action spaces, simulation or real-building validation, and explicit handling of failure modes; weaker work often relies on toy scenarios, unverified natural-language outputs, or evaluation metrics that do not translate to production building operations. This review matters because it provides a common taxonomy and maturity assessment for a rapidly fragmenting literature, helping researchers and building engineers separate near-term deployment candidates from exploratory prototypes. It also highlights the missing infrastructure needed to move LLM-assisted HVAC operations into practice, including standardized datasets, safety guardrails, benchmark suites, explainable action policies, and clear guidance on when human-in-the-loop control is required.