Aldimiro Paixão Domingos, Heleno Bispo, Thiago V. Costa, Sérgio Mauro da Silva Neiro
Abstract Crude oil refining is a complex process requiring precise modelling to optimize yield, quality, and efficiency. This study integrates Aspen HYSYS® simulations with machine learning techniques to develop predictive models for key refinery variables. Using crude oils from Arabian Light, Arabian Medium, and Gimboa, a dataset of 154 entries was generated to analyze the effects of feed composition and distillation unit operating temperatures. Four supervised learning models, linear regression (LR), random forest regression (RFR), support vector regression (SVR), and artificial neural networks (ANN), were trained and evaluated. ANN outperformed the others, achieving an R 2 score of 0.9742 and the lowest mean squared error of 0.1367, demonstrating its effectiveness in capturing nonlinear relationships. Results showed that increasing feed temperature enhances kerosene and gasoil yields while reducing the residue fraction. The study highlights the potential of combining data‐driven models with process simulations to improve refinery efficiency and decision‐making.