科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Journal of surgical education2026-09-18

Pilot Validation of a Large Language Model Facilitator for Peer-to-Peer Learning.

Simon Hollingsworth, Tom McIntyre, Sami Ahmed, Muhammad Tahir, Gavin O'Duffy, Paul F Ridgway

一句话结论 · In one sentence

This study validates the technical feasibility of LLMs in P2PL within undergraduate surgical education. However, our findings suggest LLMs are not currently capable of replacing human educators, rather they should be used as valuable tools to improve the educational experience for students. This study acts as a framework for further research in this area, ultimately paving the way for the development of a personalised LLM tutor for students.

原始摘要(英文原文)· Original abstract
OBJECTIVE: Active learning strategies such as peer-to-peer learning (P2PL) are being adopted by medical schools worldwide. P2PL involves students of the same level teaching each other without direct instructor involvement. Maintaining accuracy and consistency of information presented during P2PL necessitates employment of expert facilitators. However, additional challenges exist, such as increased costs, differences in learning opportunities, and loss of the key benefits of P2PL due to inter-facilitator differences. We aim to address these challenges by investigating three LLMs as facilitators for P2PL in surgical education amongst final year medical students. DESIGN: 65 final year medical students were recruited. P2PL presentations were scored in real-time by surgical tutors and qualitative feedback provided. P2PL presentations were recorded, transcribed, anonymised, and uploaded to three LLMs, ChatGPT, Gemini, and Claude. LLMs were prompted to score student presentations across ten metrics; identify low-confidence statements and factual errors; and provide qualitative feedback. Agreement rates between LLM and surgical tutor scoring and inter-rater reliability (IRR) was calculated. RESULTS: There was no statistically significant difference in mean score between tutors and Chat GPT and Claude (p>0.05), but a statistically significant increased level of scoring seen with Gemini (p<0.0001). IRR was assessed using intraclass correlation co-efficient (ICC). Poor reliability was seen for mean student score (ICC=0.39), moderate reliability for four scoring metrics (ICC≥0.50-<0.70) and poor reliability for six metrics (ICC<0.50). Analysis of LLM ability to identify factual errors and low confidence statements revealed poor IRR (Fleiss κ<0.20). CONCLUSIONS: This study validates the technical feasibility of LLMs in P2PL within undergraduate surgical education. However, our findings suggest LLMs are not currently capable of replacing human educators, rather they should be used as valuable tools to improve the educational experience for students. This study acts as a framework for further research in this area, ultimately paving the way for the development of a personalised LLM tutor for students.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Pilot Validation of a Large Language Model Facilitator for Peer-to-Peer Learning. — 科研速览 Science Skim