Gabriel Kadri, Raphael I. M. Santos, Felipe O. Aguiar, Victor H. O. Otani, Ricardo R. Uchida, Lucas M. Marques
Introduction Competitive performance in electronic sports (e-sports) depends on rapid decision-making, emotional regulation, and coordinated communication, yet little is known about whether pre-match vocal behavior contains information predictive of competitive outcomes. Methods This study applied supervised machine learning to 68 acoustic features extracted at the frame level (50-ms frames, 25-ms step) from 60-second pre-match team communication recordings in 89 professional Counter-Strike: Global Offensive matches; frame-level predictions were aggregated into a match-level score representing the proportion of frames classified as a win. Three predictive conditions were evaluated, acoustic features only, ranking difference only, and a combined model integrating both, using stratified group five-fold cross-validation. Uncertainty was quantified using percentile 95% confidence intervals from 2,000 match-level bootstrap resamples of the pooled out-of-fold predictions, with chance-level discrimination defined as AUC = 0.50. Results Across algorithms, models combining voice and ranking achieved the strongest performance, with the Decision Tree classifier reaching a mean bootstrap AUC of 77.3% (95% CI 65.8–86.9) and accuracy of 78.6% (95% CI 69.9–86.7); all five voice-plus-ranking models had 95% CIs excluding chance. Voice-only models showed more limited evidence of above-chance discrimination: only the Decision Tree (AUC 67.8%, 95% CI 54.8–79.2) and Random Forest (AUC 64.0%, 95% CI 50.7–76.3) had confidence intervals excluding 0.50, whereas Linear Discriminant Analysis, Logistic Regression, and k-Nearest Neighbors did not. No ranking-only model showed a confidence interval excluding chance. Exploratory LIME-based feature-attribution analyses indicated that ranking difference received the highest within-model attribution in the combined models, while delta spectral flux, chroma standard deviation, and spectral centroid received the highest within-model attribution among acoustic descriptors for tree-based, linear, and distance-based classifiers, respectively; these rankings are descriptive and were not subjected to formal cross-model statistical comparison. Discussion These findings provide preliminary, dataset-bounded evidence that acoustic patterns in brief pre-match team communication were associated with match outcome and, for some models, contributed predictive information beyond ranking; the retrospective, single-team design does not establish a generalizable behavioral biomarker or a causal link between vocal acoustics and competitive readiness.