Ahmad Shalaldeh, Majeed Safa, Chris Logan, Mohmmad Othman
The non-invasive determination of live weight and body composition of ewes is an important element in ensuring precision livestock management and animal well-being. Traditional practices tend to be subjective, labor-intensive, or rely on expensive medical imaging such as Computed Tomography (CT). This paper proposes a new hybrid deep learning method to predict live weight and carcass traits in Coopworth ewes. The dataset of 1184 images taken from 156 ewes was analyzed and compared using a hybrid model (ResNet18 with Multi-Layer Perceptron through simple concatenation) and two more advanced models: Attention-Guided Feature Fusion Network (AGFF-Net) based on cross-modal attention and a Vision Transformer-based Hybrid Regressor (ViT-HR). Auxiliary tabular variables are the Body Condition Score (BCS) and size category. The Transformer architecture predicts (R2 = 0.93) the live weight of ewes by dynamically ranking each visual patch and asking it to query the self-attention sequence. This technique treats the BCS as a distinct token in the self-attention sequence. Data partitioning at the animal level was stringent, thereby giving strong generalization. Findings indicate that the best advanced fusion systems are far better than baseline concatenation, with a high accuracy confirmed with gold standards obtained by CT. Grad-CAM visual explainability makes sure that models are able to localize biologically relevant anatomical locations successfully. The study closes the gap between complex deep learning models and real-world agriculture implementation to provide a correct, interpretable and scalable solution to real-time livestock measurements.