Safe, Adaptive, and Scalable AI-Driven Energy Management of Urban Multi-Energy Systems
Loading...
Date
Authors
Journal Title
Journal ISSN
Volume Title
Publisher
University of Waterloo
Abstract
Cities serve as central hubs of economic, industrial, and social activity across nations. However, the high population densities, essential infrastructure networks, and diverse industrial operations position cities as the largest energy consumers and the primary contributors to greenhouse gas (GHG) emissions. This poses critical energy and environmental challenges on the urban scale. These challenges include the need to enhance energy generation and consumption efficiency and to achieve a substantial reduction in GHG emissions across different energy sectors. One promising solution is the transition from the traditional, independent planning and operation of energy sectors to the integrated planning and operation of Urban Multi-Energy Systems (UMESs). In this context, a UMES can be defined as a city-scale energy system that coordinates electricity, natural gas, heating, cooling, and other energy sectors as a single integrated system. This integration begins at the multi-energy buildings level, expands to community-level Multi-Energy Microgrids (MEMGs), and ultimately forms city-scale Networked Multi-Energy Microgrids (N-MEMGs) connected through electricity, gas, and heat distribution networks. By coordinating multiple energy carriers, UMESs can facilitate integration of variable renewable energy, improve overall energy efficiency, and enhance operational flexibility and resilience.
Despite these advantages, energy management of UMESs is a particularly challenging operational problem. This is due to complex multi-energy flows, uncertainty propagation,
decentralized ownership, privacy limitations, and the need for emission-aware operation. In this context, the energy management of UMESs has been addressed using two distinct approaches: model-driven and AI-driven. Model-driven approaches can explicitly represent physical constraints and yield interpretable solutions; however, they often suffer from high computational complexity and rely on complex mathematical optimization models,
thereby limiting their suitability for real-time applications. On the other hand, AI-driven approaches offer adaptive decision-making capabilities under uncertainty. Among these approaches, Deep Reinforcement Learning (DRL), which formulates sequential optimization problems as Markov Decision Processes (MDPs), has gained significant attention. However, conventional DRL methods lack explicit mechanisms to enforce physical constraints, thereby limiting their applicability to real-world energy management systems.
The primary objective of this thesis is to develop a safe, adaptive, and scalable AI-based energy management framework for UMESs. The proposed framework models the system as a Constrained Markov Decision Process (CMDP) and employs Constrained Deep Reinforcement Learning (CDRL) algorithms to solve it. The framework aims to generate low-carbon and economically efficient dispatch strategies while adhering to the operational and engineering constraints inherent in physical multi-energy networks. Additionally, it accounts for uncertainties related to renewable energy resources, energy prices, and energy demand.
The first part of the thesis addresses the energy management problem of single MEMG systems through a hierarchical two-layer framework. In the first layer, a model-driven Mixed-Integer Quadratic Programming (MIQP) formulation performs day-ahead scheduling by minimizing operational costs subject to the relevant system constraints. In the second layer, a CDRL agent performs real-time adjustments to compensate for uncertainty and operational deviations. This framework combines the strength of CDRL in sequential decision-making under stochastic conditions with the structural and feasibility guarantees offered by model-driven mathematical programming. In this way, the day-ahead optimization layer guides the agent’s exploration during training and reduces reliance on black-box policies during real-time operation.
Although the model-driven day-ahead layer developed in the first part provides valuable guidance for the CDRL agent’s decisions, it becomes computationally expensive as system complexity increases. Therefore, the second part of the thesis eliminates this dependence and solves the energy management problem for a single MEMG using a fully model-free CDRL approach. This approach incorporates an integrated demand response (IDR) strategy that jointly exploits demand-side flexibility across electricity, heat, and gas, and it models carbon capture, storage, and utilization (CCSU) facilities that enable participation in carbon trading markets (CTMs). The operational dynamics of the system
are formulated as CMDP and solved using the interior-point policy optimization (IPO) algorithm. The proposed framework yields dispatch decisions that improve both economic and environmental performance while ensuring satisfaction of system constraints.
The final part of the thesis extends the approach to cooperative energy management of networked MEMG systems using a Constrained Multi-Agent DRL (CMADRL) method. Here, the problem is formulated as a constrained Decentralized Partially Observable Markov Decision Process (Dec-POMDP). This jointly captures the interactions between each MEMG operator and the energy network operators, such as power, gas, and heat. The Dec-POMDP
balances a cooperative economic-emission objective with physical network constraints. This formulation explicitly accounts for both the internal distributed generators (DGs) within each MEMG and the cross-network coupling DGs. It further defines the necessary components for multi-agent training, including decentralized observations, agent-specific action spaces, and a cooperative reward. In addition, a novel Sensitivity-Weighted Constraint Decomposition (SWCD) mechanism, derived from network sensitivities, is developed to compute per-agent physics-informed constraint-cost signals, enabling the training of multiple agents for safe and cooperative dispatch. CMADRL is trained using the HeterogeneousAgent Proximal Policy Optimization-Lagrangian (HAPPO-Lag) algorithm. The resulting physics-informed framework enforces constraints across the coupled power, gas, and heat distribution networks in a decentralized and privacy-preserving manner.
Collectively, these three contributions establish a unified, safe, adaptive, and scalable AI-driven framework that progresses from single MEMG to city-scale networked multienergy systems. The proposed methods deliver low-carbon, economically efficient energy dispatch while ensuring compliance with the physical and engineering constraints governing coupled power, gas, and heat networks. The framework thus offers a practical pathway
toward the real-time, privacy-preserving, and emission-aware energy management of future urban multi-energy systems.