From Digital Twin to Real Building: Training Adaptive Agents for HVAC Control
Abstract
Buildings account for around 30\% of global energy demand, making more efficient HVAC operation an important emissions-reduction opportunity. While Reinforcement Learning (RL) promises substantial efficiency gains, its real-world climate impact is constrained by two challenges: operational trust and modeling scalability. We present a practical, three-stage sim-to-real framework that directly resolves both bottlenecks. To unlock scalability, we leverage a lightweight, floorplan-based digital twin calibrated rapidly from telemetry. To establish trust, we initialize the policy via historical behavior cloning, audit learned strategies with attribution methods to verify anticipatory control, and enforce deterministic guardrails with provable compliance bounds during physical deployment. In simulation, our selected policy achieves a 79.9% reduction in the normalized energy metric while maintaining comfort. A live commercial office deployment confirms closed-loop 24/7 execution across coupled airside and waterside systems with progressive online adaptation. This framework provides a trusted, scalable blueprint for accelerating operational emissions reductions across the existing building stock.