Yogesh Kumar, Inderpreet Kaur, Nandini Modi, Priya Bhardwaj, Jaeyoung Choi, Muhammad Fazal Ijaz
The highest held-out test accuracy (96.87%) was obtained with the Custom CNN from among the configurations that were saved. The MLP Mixer was the most highly accurate model during training (99.90%), but the model did not perform as well on the validation set (93.33%) or held-out test set (93.73%), suggesting some in-sample fitting. Model-capacity information, parameter-to-sample ratios, learning curves, class-wise metrics, and a targeted class-weighting comparison are reported. Class-weighting was performed during the training phase only, but class-weighting did not yield highest performance for the tested Custom CNN configuration. of the chosen Custom CNN configuration for the observed dataset.
INTRODUCTION: Computational classifiers can be benchmarked against public encoded splice-junction datasets, but the performance estimate will depend on the input representation, preprocessing pipeline, model set-up and evaluation design.
METHODS: This study investigates the performance of 15 different Convolutional, Hybrid and Deep Learning architectures on a public encoded splice-junction dataset for 3-class classification. The results of the primary models were archived from one train/validation/held-out-test workflow and are presented as illustrative effects and not as statistically significant evidence of superiority.
RESULTS: The highest held-out test accuracy (96.87%) was obtained with the Custom CNN from among the configurations that were saved. The MLP Mixer was the most highly accurate model during training (99.90%), but the model did not perform as well on the validation set (93.33%) or held-out test set (93.73%), suggesting some in-sample fitting. Model-capacity information, parameter-to-sample ratios, learning curves, class-wise metrics, and a targeted class-weighting comparison are reported. Class-weighting was performed during the training phase only, but class-weighting did not yield highest performance for the tested Custom CNN configuration. of the chosen Custom CNN configuration for the observed dataset.
DISCUSSION: The conclusions are limited to the classification results that could be computed for the encoded benchmark dataset and the reported evaluation procedure. The results do not demonstrate a statistically significant superiority of the models in the sense of external generalization, clinical diagnostic validity, patient-based mutation detection, clinical utility or readiness for deployment. Leakage-free repeated evaluation, external datasets, stored sample predictions, statistically comparison of models' outputs, and clinically validated patient-level data would be needed for more general conclusions.