Antonis Zakynthinos, Vasilis Michalakopoulos, Elissaios Sarmas, Vangelis Marinakis
The ongoing evolution of energy systems is shifting conventional power grids toward smart grids, where advanced metering infrastructure and smart meters provide detailed building-level data. However, many buildings (residential or not) are characterized by historical load data scarcity, which constrains the accuracy of short-term load forecasting models, necessary for participating in energy management schemes. Transfer Learning (TL) provides a compelling approach by enabling pre-trained models to be adapted to data scarce environments, effectively leveraging knowledge acquired from data-rich domains. This study explores a systematic experimental framework consisting of four established baseline models alongside the state-of-the-art Temporal Fusion Transformer. Each model is pretrained on a ”source” building (data-rich) and subsequently transferred to target (data-poor) buildings of varying domain similarity, simulating scenarios of data scarcity. To enable effective knowledge transfer, three to five parameter-based transfer learning techniques were tested per baseline model, including full fine-tuning and partial layer adaptation, while a total of three novel fine-grained knowledge transfer methods were applied on the Temporal Fusion Transformer targeting its variable selection, recurrent and attention components. Experimental results on a real-world dataset containing student dormitory buildings from three different geographical locations, demonstrate that TL consistently enhances forecasting accuracy across all target domains, with the largest benefits observed under severe data scarcity, where the best performing TL configurations achieved an RMSE reduced by 10-20% compared to base models. Even in settings with larger domain gaps or longer historical records, TL still yielded stable improvements in all metrics across all models while the Temporal Fusion Transformer notably yielded around 20% reduced RMSE and 41% increased R 2 Score. Moreover, it surpassed the strongest TL baselines, delivering up to 22% additional improvement in R 2 Score and consistent reductions in RMSE and MAE across domains. These findings highlight both the effectiveness of TL in mitigating data scarcity and the robustness of Temporal Fusion Transformer in extracting temporal and contextual patterns, confirming its value for cross-domain and limited-data load forecasting.