科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Frontiers in physiology2026-01-01· Lesson plan

Physiological relevance and generation-level stability of artificial intelligence-generated race-day warm-up and cool-down plans for sprint swimming: an expert-rated comparison of four large language models.

TianYue Zhao, YiYao Peng, ZiXin Chen, Xiaojun Wang

一句话结论 · In one sentence

Publicly accessible LLM-based AI tools can generate structured swimming race-day warm-up and immediate post-race cool-down plans, but their ability to translate sprint-swimming physiological principles into practical prescriptions differs across tools and plan sections. In this simulated 100-m freestyle scenario, ChatGPT showed the most favorable overall expert-rated profile, whereas Doubao produced more stable but lower-rated outputs. These findings suggest that AI-generated swimming race-day plans should not be used as stand-alone prescriptions. Instead, they may support coaches by providing preliminary drafts, alternative plan structures, and clearer prescription details for further expert screening and revision. Final decisions on intensity, sequencing, recovery, risk management, and individual adaptation should remain under qualified coach supervision.

原始摘要(英文原文)· Original abstract
BACKGROUND: Large language model (LLM)-based artificial intelligence (AI) tools are increasingly used to generate sport and exercise training content. However, the content quality, sport-specific relevance, safety, and repeated-generation stability of AI-generated race-day plans remain insufficiently examined, particularly in competitive swimming. It also remains unclear how different AI tools may support coaches in drafting, comparing, and refining competition-day preparation plans. OBJECTIVE: This study aimed to compare the expert-rated physiological relevance, content quality, and generation-level stability of race-day warm-up and immediate post-race cool-down plans generated by four publicly accessible LLM-based AI tools for a simulated 100-m freestyle swimmer, and to identify how these tools may serve as coach-supervised drafting and comparison aids in swimming race-day preparation. METHODS: A simulation-based content-analysis design with a within-rater repeated-measures structure was used. Four AI tools-ChatGPT, Gemini, DeepSeek, and Doubao-each generated five independent outputs using the same standardized detailed prompt. Each output included a race-day warm-up plan and an immediate post-race cool-down plan. After anonymization and randomization, 32 experts with experience in coaching or teaching 100-m sprint swimming independently rated the plans using a 40-item multidimensional rating instrument. Inter-rater reliability was assessed using intraclass correlation coefficients and Krippendorff's alpha. The primary analysis used a cumulative link mixed model with fixed effects for AI model, section, and their interaction, and random intercepts for rater, item, and generation ID. Exploratory domain-level and generation-level stability analyses were also conducted. RESULTS: ChatGPT showed the highest descriptive overall quality. The cumulative link mixed model showed a significant model × section interaction (p <.001), indicating that model differences varied between the warm-up and cool-down sections. ChatGPT received higher ratings than DeepSeek and Doubao in the warm-up section and than all three other models in the cool-down section. Exploratory domain-level analyses showed differences across all 10 content-quality domains. Generation-level stability did not correspond directly to expert-rated quality: Doubao was the most stable model, whereas ChatGPT showed the most favorable overall quality profile but greater repeated-generation variability. CONCLUSIONS: Publicly accessible LLM-based AI tools can generate structured swimming race-day warm-up and immediate post-race cool-down plans, but their ability to translate sprint-swimming physiological principles into practical prescriptions differs across tools and plan sections. In this simulated 100-m freestyle scenario, ChatGPT showed the most favorable overall expert-rated profile, whereas Doubao produced more stable but lower-rated outputs. These findings suggest that AI-generated swimming race-day plans should not be used as stand-alone prescriptions. Instead, they may support coaches by providing preliminary drafts, alternative plan structures, and clearer prescription details for further expert screening and revision. Final decisions on intensity, sequencing, recovery, risk management, and individual adaptation should remain under qualified coach supervision.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Physiological relevance and generation-level stability of artificial intelligence-generated race-day warm-up and cool-down plans for sprint swimming: an expert-rated comparison of four large language models. — 科研速览 Science Skim