Rafeef Fauzi Najim Alshammari, Zainab Khudhur Mohsin, Wamedh Hazem Sahib Al-Tufaili
Facial emotion recognition (FER) has moved from a niche research topic into a standard component of affective computing, yet the ensemble models that achieve the best accuracy tend to be too heavy for devices at the edge.This paper proposes a multi-level knowledge distillation (KD) framework that compresses EA-Net an ensemble of EfficientNet-B0 and InceptionV3 augmented with channel (CAM) and spatial (SAM) attention into three lightweight students: EfficientNet-B0, MobileNetV3-Small with lightweight attention, and ShuffleNetV2 0.5× with lightweight attention.Knowledge transfer happens at three representation levels through a progressive curriculum that layers cross-entropy supervision, temperature-scaled logit distillation, intermediate feature matching, and attention map transfer across successive stages.All models are evaluated on FER2013 and KDEF with a reproducible pipeline (deterministic seeding, JSONL logging, automated figure generation).The ShuffleNet student is the strongest of the three on FER2013 and the nominal best on KDEF (69.8% accuracy, 0.670 macro-F1 on FER2013; 81.8% accuracy, 0.819 macro-F1 on identity-disjoint KDEF), while remaining compact.MobileNetV3-Small ranks second on FER2013 (on KDEF the three students fall within one standard deviation of one another and are not reliably ranked) and offers the lowest CPU and GPU forward-pass latency, which makes it attractive when latency is the binding deployment constraint.Grad-CAM visualisations confirm that the distilled students align their spatial focus with the teacher's attention maps on the mouth and eye regions, supporting the use of multi-level KD for compact yet interpretable FER models.