Mediana Aryuni, Chastine Fatichah, Anny Yuniarti
Multi-label Classification (MLC) has become crucial in various real-world applications, but its performance is often influenced by complex data imbalance issues, inter-label dependencies, and label concurrency.The imbalance in MLC is much more complicated than in single-label classification because minority labels often appear alongside majority labels, leading to model bias.This research proposes
Multi-label Classification (MLC) has become crucial in various real-world applications, but its performance is often influenced by complex data imbalance issues, inter-label dependencies, and label concurrency.The imbalance in MLC is much more complicated than in single-label classification because minority labels often appear alongside majority labels, leading to model bias.This research proposes an extended hybrid resampling architecture combined with a Heterogeneous Stacking Ensemble to address these issues.This methodology begins with a two-stage hybrid resampling technique: first, using the R-HwR-ROS and R-HwR-SMT for label decoupling and dataset reconstruction to reduce concurrency.Second, applying the Multi-Label edited Nearest Neighbor (MLeNN) module for noise filtering and Multi-Label Tomek Link (MLTL) for borderline refinement.In the classification stage, three heterogeneous paradigms are used as base learners: Classifier Chain (CC), Calibrated Label Ranking (CLR), and ML-kNN.The predictions from these three models are then combined using XGBoost as a meta-learner optimized to handle the complexity of inter-label relationships.Evaluation uses ten predefined outer folds, five-fold OOF probability generation, and training-only threshold selection that maximizes F1-Macro.Across eight datasets, the best E5 configuration achieved a higher best-stage value than E4 on five datasets for F1-Macro, six for F1-Micro, and all eight for Subset Accuracy; it achieved a lower Hamming Loss on seven datasets, but a higher Macro Recall on only one dataset.The exploratory across-dataset Wilcoxon-Holm comparison identified a significant E5 improvement only for Subset Accuracy (pHolm=0.0391);the F1-Macro difference was not significant (pHolm=0.4609).These findings show that stacking improves selected performance dimensions but does not uniformly improve minority-label sensitivity, so conclusions must combine F1-Macro, Macro Recall, per-label behavior, and paired statistical evidence.