Víctor-Manuel Vargas, David Guijo-Rubio, Rafael Ayllón-Gavilán, Antonio M. Gómez-Orellana, Pedro Antonio Gutiérrez, and César Hervás-Martínez
Ordinal classification, where labels follow a natural order, has gained increasing attention, particularly in the deep learning community due to its relevance in tasks such as age estimation, medical grading, and quality assessment. Despite the growing number of deep ordinal classification methods, a comprehensive experimental analysis of their core ordinal components remains lacking. This work presents a systematic evaluation of deep ordinal classifiers by analysing the impact of three key modelling choices: the loss function, output layer, and labelling strategy. To analyse their effects, we adopt a unified architecture and evaluate one nominal and 19 ordinal configurations, resulting from combination of two loss functions, two output layers, and five labelling strategies. These configurations are assessed on 12 diverse ordinal image datasets using six performance metrics, including both ordinal and nominal measures. Results show that ordinal output layers consistently outperform softmax, and that soft labelling generally improves generalisation. While categorical cross-entropy achieves better average performance, especially on nominal metrics, no configuration performs best across all datasets. Statistical analyses indicate significant interactions between losses, outputs, labelling strategies, and datasets, highlighting the need to adapt methodological choices to specific tasks. These findings provide valuable guidance for designing robust deep ordinal classification models.