科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Frontiers in artificial intelligence2026-01-01

Evaluating the robustness of specialized and general-purpose facial expression recognition systems across varied scenarios.

José Salas-Cáceres, Javier Lorenzo-Navarro, Modesto Castrillón-Santana, Patricia Picazo-Peral, Sergio Moreno-Gil

一句话结论 · In one sentence

The results show that performance on controlled datasets substantially overestimates real-world FER capability, with average weighted and unweighted average recall values decreasing from approximately 72% in static datasets to below 30% in naturalistic settings. All tested models exhibited a marked bias toward happiness, with negative emotions frequently misclassified, a trend particularly pronounced in VLMs, where categories such as fear or anger often received F1-scores near zero. Among the neural networks, the DAN model trained on the AfectNet dataset achieved the strongest generalization, outperforming all VLMs and confirming that AfectNet provides a more realistic training distribution than the RAF-DB database. FaceReader© delivered excellent performance under ideal conditions but experienced substantial degradation in dynamic scenarios, falling below a random classifier in CREMA-D.

原始摘要(英文原文)· Original abstract
INTRODUCTION: This work presents a comprehensive evaluation of facial expression recognition (FER) systems across four benchmarked datasets of varying complexity, ranging from controlled static images (ADFES, WSEFEP) to more realistic dynamic recordings (RAVDESS, CREMA-D). METHODS: Three categories of models were evaluated: traditional FER neural networks models, general-purpose vision language models (VLMs), and the commercial software FaceReader© 10. RESULTS: The results show that performance on controlled datasets substantially overestimates real-world FER capability, with average weighted and unweighted average recall values decreasing from approximately 72% in static datasets to below 30% in naturalistic settings. All tested models exhibited a marked bias toward happiness, with negative emotions frequently misclassified, a trend particularly pronounced in VLMs, where categories such as fear or anger often received F1-scores near zero. Among the neural networks, the DAN model trained on the AfectNet dataset achieved the strongest generalization, outperforming all VLMs and confirming that AfectNet provides a more realistic training distribution than the RAF-DB database. FaceReader© delivered excellent performance under ideal conditions but experienced substantial degradation in dynamic scenarios, falling below a random classifier in CREMA-D. DISCUSSION: These findings highlight the limitations of general-purpose VLMs and commercial tools for real-world FER and underscore the need for models explicitly designed to handle naturalistic variability. Furthermore, the reported performance of FaceReader© 10 in their manual on ADFES and WSEFEP was corroborated in this study.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Evaluating the robustness of specialized and general-purpose facial expression recognition systems across varied scenarios. — 科研速览 Science Skim