Auditable Reinforcement Learning for Energy-Efficient Building HVAC Control
Abstract
Heating, ventilation, and air-conditioning (HVAC) accounts for approximately half of building energy demand, so how it is controlled is a direct lever on building emissions. Since setpoint changes have delayed effects on zone temperatures, Reinforcement learning (RL) is well suited to HVAC control, and RL estimates the long-run value of an action directly from observed transitions. A conventional neural RL controller, however, does not provide a direct way to reveal why a setpoint is selected based on the building state or how it is compared with alternative setpoints. To make learned HVAC controller auditable, we propose Orthogonal Adaptive Rule Q-learning (OARQ), which retains fitted Q-learning but represents the Q-value of each candidate setpoint as a compact additive ensemble of intrinsically interpretable rules. The rules are learned from data rather than written by an engineer or extracted from a trained network afterwards. The ensemble is additive over conditions on named building variables, so the rules that fire and their weights are directly readable. It is the deployed controller rather than a model of it, so they account for the setpoint and its margin exactly. In a 19-zone EnergyPlus office evaluated across four climate--season settings, OARQ improves zone comfort over a tuned neural baseline while matching or reducing HVAC electricity. It therefore attains competitive control performance while remaining directly auditable, so an operator can check an energy-saving setpoint against the conditions.