Reinforcement Learning with Verifiable Rewards for HVAC Control
Takumi Shioda ⋅ Tatsuo Nagai
Abstract
Buildings account for about 30% of global final energy consumption, and heating, ventilation and air conditioning (HVAC) systems account for almost half of energy use in buildings. Advanced controllers can reduce this energy use, but existing approaches are difficult to scale because they usually require site-specific models, data, or training. We post-train an open-weight large language model once with reinforcement learning from verifiable rewards (RLVR) to learn a reusable decision procedure, which we then apply to new sites using only text inputs. We use thermal energy storage planning under time-varying grid carbon intensity as our test task. Because the value of an HVAC action taken now depends on future operating conditions, we compute optimal long-horizon action values with dynamic programming and use them as dense verifier rewards. Using only 30 prompts from one site, RLVR closes 96% of the base model's gap to the DP optimum. Without further weight updates, it closes 82% of the gap at a site with different equipment sizes and carbon-intensity profile, reducing CO$_2$ emissions by 10% relative to the base model. These results suggest that RLVR offers a scalable path to carbon-aware HVAC operation without site-specific retraining.
Chat is not available.
Successful Page Load