Abhay Bhandarkar, S. Sridhar, N. Kirn Kumar, Madhumitha R., Rahul S.G., Amirthalakshmi T.M.
Accurate short-term wind power forecasting is critical for grid stability and renewable energy integration, yet existing deep learning approaches treat all input features uniformly, ignoring known physical subsystem structure, and require long historical windows to achieve competitive accuracy. This study introduces a Physics-Informed Temporal Fusion Network (PI-TFN) that bridges the gap between domain-specific physical knowledge and deep temporal learning through architectural inductive bias. The proposed architecture decomposes wind turbine dynamics into three specialized processing pathways: a physics-informed layer encoding aerodynamic relationships (including the cubic wind speed–power law and a learned power coefficient estimator), a thermal state encoder capturing temperature dynamics via a GRU, and a blade coordination module modeling pitch angle synchronization. These are fused with a bidirectional LSTM backbone augmented by multi-scale temporal attention. A composite training objective combines mean squared error with soft physical penalties for cut-in, rated-power saturation, non-negativity, and a power-curve residual that ties predictions to the theoretical aerodynamic relationship. All models are trained under a strictly causal preprocessing pipeline in which missing values are imputed by forward fill only, eliminating any leakage of future information into past observations. Evaluated on 118,224 SCADA observations spanning 28 months, PI-TFN achieves the lowest RMSE among five deep learning baselines (LSTM, TCN, Transformer, TFT, Informer) at a four-hour lookback window. Multi-step experiments show that the advantage of the physics-guided design grows with the forecast horizon: at six hours ahead, PI-TFN attains 535.3 kW RMSE against 607.9 kW for LSTM and 637.3 kW for the persistence baseline, the latter of which can no longer track the wind generation process at that horizon. A capacity-controlled comparison, in which a 348,114-parameter PI-TFN variant outperforms a 376,061-parameter TFT and a 601,601-parameter LSTM, confirms that the improvements stem from the physics-guided decomposition rather than from raw model capacity. External validation on the Kelmarsh wind farm SCADA dataset (six Senvion MM92 turbines pooled, 213,050 training samples) and on the Open Power System Data German national wind generation series tests how the proposed design transfers across geographies, temporal resolutions, and aggregation levels. A sample-efficiency experiment on Kelmarsh shows that PI-TFN’s relative advantage over an LSTM baseline is largest in the data-scarce regime (15.9 kW better RMSE at 5000 training samples) and reverses as the training set grows, confirming that the physics-informed pathways supply inductive bias most usefully where data is limited.