π€ AI Summary
This study addresses the issue of excessive compressor cycling in residential heat pumps, which accelerates equipment wearβa factor commonly overlooked by existing reinforcement learning controllers that focus primarily on energy consumption and thermal comfort. To bridge this gap, the work explicitly incorporates compressor wear into the reward function and evaluates Soft Actor-Critic (SAC) and Proximal Policy Optimization (PPO) algorithms within the hydronic heat pump case of the BOPTEST platform. Results demonstrate that SAC autonomously learns a variable-speed continuous modulation strategy, achieving zero start-stop cycles while reducing thermal discomfort by 90.7% at the cost of only an 11.5% increase in operational expenditure, substantially outperforming conventional control approaches.
π Abstract
On--off cycling is the main cause of compressor wear in residential heat pumps, yet reinforcement learning (RL) controllers for buildings typically optimise only energy cost and thermal comfort, ignoring how much the learned policy cycles. We add a levelised compressor-wear term to the control reward and study how the resulting behaviour depends on the RL algorithm. Training Soft Actor---Critic (SAC) and Proximal Policy Optimisation (PPO) on an identical Markov decision process for the BOPTEST bestest hydronic heat pump case, we find that SAC learns a continuous modulation policy that keeps the compressor permanently engaged---the operating principle of an inverter-driven heat pump---achieving zero start-ups per day, whereas PPO collapses to bang-bang control that cycles more than the baseline. On the BOPTEST emulator the SAC policy cuts thermal discomfort by up to 90.7% for an 11.5% cost increase, while eliminating all baseline cycling.