科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Frontiers in artificial intelligence2026-01-01

A unified MAP-EM approach to stable Gaussian mixture clustering with priors, graphs, and split-merge adaptation for document clustering.

Sumathi Subbarayan, G Hannah Grace

一句话结论 · In one sentence

On Reuters-R8, the model achieved an Adjusted Rand Index of 0.347 and an accuracy of 0.569, outperforming the variational Bayesian Gaussian mixture model (GMM) (0.317). On the BBC dataset, it achieved the highest Adjusted Rand Index of 0.326 and a Normalized Mutual Information of 0.425 among the methods compared. Formal statistical testing showed that MAP-EM achieved significant positive differences in 45 out of 75 method-level comparisons, with one significant negative comparison. At the run level, MAP-EM obtained higher scores in 891 out of 1,115 valid paired comparisons, corresponding to a win rate of 79.9%. Theoretical analysis further supports the proposed framework by establishing coercivity of the penalized objective, monotone ascent of the MAP-EM iterations, finite termination of the split-merge stage under the stated acceptance criterion, boundedness of the covariance estimates under the normal-inverse-Wishart prior, local R-linear convergence of the MAP-EM iterations after model-order stabilization, eigenvalue bounds for the NIW covariance estimator, a condition-number bound for the penalized mean update, and a perturbation bound for the graph-regularized mean update.

原始摘要(英文原文)· Original abstract
INTRODUCTION: Clustering high-dimensional and noisy data remains challenging for conventional expectation-maximization (EM) methods as overlapping clusters, sparse features, and outliers can lead to covariance degeneracy and unstable parameter estimates. This research aims to improve clustering performance in high-dimensional, noisy settings by developing a robust maximum a posteriori expectation-maximization (MAP-EM) framework that integrates prior regularization, geometric structure, and outlier handling. Traditional EM-based clustering methods often struggle in the presence of overlapping clusters, high-dimensional features, and outliers, leading to degenerate covariance and unstable parameter estimates. METHODS: The proposed MAP-EM model improves reliability by combining normal-inverse-Wishart (NIW) priors for covariance stabilization, a graph-Laplacian structure over the feature space to capture geometric relations among features, a uniform noise component to absorb outliers, and an adaptive split-merge strategy that refines cluster boundaries. These modules are coupled within a single MAP-EM procedure. The noise component modifies the E-step responsibilities by capturing atypical observations, and these updated responsibilities drive the NIW-regularized covariance and the graph-regularized mean updates in the M-step. The split-merge step is accepted only if it improves the penalized objective. The proposed MAP-EM model was evaluated on five synthetic datasets and three benchmark text corpora, namely Reuters-R8, BBC Sports, and the BBC dataset. RESULTS: On Reuters-R8, the model achieved an Adjusted Rand Index of 0.347 and an accuracy of 0.569, outperforming the variational Bayesian Gaussian mixture model (GMM) (0.317). On the BBC dataset, it achieved the highest Adjusted Rand Index of 0.326 and a Normalized Mutual Information of 0.425 among the methods compared. Formal statistical testing showed that MAP-EM achieved significant positive differences in 45 out of 75 method-level comparisons, with one significant negative comparison. At the run level, MAP-EM obtained higher scores in 891 out of 1,115 valid paired comparisons, corresponding to a win rate of 79.9%. Theoretical analysis further supports the proposed framework by establishing coercivity of the penalized objective, monotone ascent of the MAP-EM iterations, finite termination of the split-merge stage under the stated acceptance criterion, boundedness of the covariance estimates under the normal-inverse-Wishart prior, local R-linear convergence of the MAP-EM iterations after model-order stabilization, eigenvalue bounds for the NIW covariance estimator, a condition-number bound for the penalized mean update, and a perturbation bound for the graph-regularized mean update. DISCUSSION: The proposed MAP-EM framework provides a stable, structure-aware clustering approach for high-dimensional text data. Experimental results indicate that its advantages are most evident on datasets with noise, overlapping clusters, and well-connected feature graphs, rather than across all clustering scenarios.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

A unified MAP-EM approach to stable Gaussian mixture clustering with priors, graphs, and split-merge adaptation for document clustering. — 科研速览 Science Skim