Jun Tamura, Yuki Itaya, Kenichi Hayashi, Kouji Yamamoto
Classification problems underpin critical decisions, such as predicting patient prognosis and selecting treatment strategies. Accordingly, accurate evaluation of classification models is essential, and numerous metrics have been proposed. Among them, the Matthews correlation coefficient (MCC)–also called the phi coefficient–offers a balanced assessment even under severe class imbalance. However, the growing prevalence of multiclass problems (three or more classes) has led researchers to adopt macro-averaged and micro-averaged versions of MCC without formal definitions or authoritative references. This study supplies that foundation by formally defining both extensions and detailing their statistical properties. To move beyond point estimates, we develop several methods for constructing asymptotic confidence intervals for the proposed metrics. We further extend these techniques to derive confidence intervals for paired differences, enabling rigorous comparison of competing classifiers within the same subjects. The performance of the intervals is investigated through comprehensive simulation studies and demonstrated on real-world biomedical datasets. Our framework equips practitioners with principled, readily implementable tools for statistically sound evaluation of multiclass classification performance.