Faizal Widya Nugraha, Thanda Shwe, Israel Mendonça, Masayoshi Aritsugi
In the data-driven applications, class overlap and class imbalance are major factors that degrade machine learning performance. While existing resampling methods often address these issues independently, they frequently fail to preserve the underlying distribution of the majority class after cleaning. This study proposes Two-Phase Clustering Plus, a novel framework that uniquely leverages overlap detection to inform the undersampling process. Unlike conventional approaches that remove samples indiscriminately, this framework integrates the Edited Nearest Neighbors technique to isolate overlapping regions, applying data cleaning strictly to the majority class. To manage the trade-off between overlap reduction and the preservation of informative boundary samples, this localized cleaning strategy is combined with a two-phase clustering strategy that retains the safe majority data. This synergy ensures the targeted elimination of ambiguous majority instances without causing excessive information loss, systematically selecting representative samples that maintain critical boundaries and the global data structure. Evaluated on five service industry datasets representing virtual sensing environments derived from customer behavioral data, experimental results demonstrate that the proposed method yields a 1.45% improvement in geometric mean and significantly enhances class separability. By explicitly linking overlap removal with strategic clustering, this approach provides a robust and distribution-aware solution for complex imbalanced datasets.