科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Frontiers in psychology2026-01-01

Revisiting judging reliability in taekwondo freestyle Poomsae: implications for AI-supported evaluation.

Min-Woo Jeon, Hong-Suk Kim, Ji-Yong Park, Hye-Soo Cho

一句话结论 · In one sentence

Systematic session effects, absolute test-retest agreement, relative consistency over time, and inter-rater agreement represent distinct aspects of judging reliability. Evaluation components with poor single-judge agreement require clearer operational definitions and further validation before computational or AI-assisted scoring is considered.

原始摘要(英文原文)· Original abstract
BACKGROUND: Before artificial intelligence-based judging systems can be applied to Taekwondo Poomsae, it is necessary to understand which performance components human judges can identify consistently and which remain difficult to evaluate reliably. This study examined reliability in freestyle Poomsae judging to identify evaluation components with stable and unstable human scoring. METHODS: Ten internationally certified Taekwondo Poomsae referees evaluated ten official competition videos twice, separated by a one-week washout period. Systematic session effects were evaluated with separate linear mixed-effects models estimated by restricted maximum likelihood using the SPSS MIXED procedure. Session was specified as a fixed effect, judge and video as crossed random intercepts, and the two observations within each judge-video pair as repeated measurements with an unstructured residual covariance matrix. Fixed-effect inference used Satterthwaite denominator degrees of freedom. Test-retest stability was evaluated for identical judge-video pairs, with primary emphasis on the absolute-agreement ICC because systematic session effects were detected. Inter-rater agreement was evaluated separately by session using two-way random effects, single rating absolute-agreement and consistency ICCs with 95% confidence intervals and Kendall's W. Item-specific analyses were exploratory, and no multiplicity adjustment was applied. RESULTS: The total score increased by 0.304 points, 95% CI [0.187, 0.421], t (99) = 5.177, p < 0.001. At the nominal 0.05 level, exploratory item-specific increases were observed for jumping side kick (β = 0.057, p < 0.001), acrobatic kicking technique (β = 0.021, p = 0.003), and expression of energy (β = 0.047, p = 0.003); the increase for basic movements and practicability did not reach the 0.05 threshold (p = 0.053). Absolute-agreement test-retest ICCs ranged from 0.039 to 0.848 and were excellent for harmony and the total score. All single-rating inter-rater ICC(A,1) point estimates were below 0.40. For the total score, ICC(A,1) was -0.007, 95% CI [-0.018, 0.036], in the first session and 0.060, 95% CI [0.010, 0.224], in the second session. CONCLUSION: Systematic session effects, absolute test-retest agreement, relative consistency over time, and inter-rater agreement represent distinct aspects of judging reliability. Evaluation components with poor single-judge agreement require clearer operational definitions and further validation before computational or AI-assisted scoring is considered.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Revisiting judging reliability in taekwondo freestyle Poomsae: implications for AI-supported evaluation. — 科研速览 Science Skim