科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Surgical endoscopy2026-09-21

SAFER: a simulation-based assessment tool for emergency robotic undocking with multisource validity evidence.

Blake T Beneville, Abigail J Hatcher, Michael M Awad

一句话结论 · In one sentence

This proof-of-concept study shows that an LLM can generate structured, rubric-mappable OSCE responses and may help educators flag missing, ambiguous, or inconsistent checklist elements for expert review. These exploratory findings support further investigation of LLMs as decision-support tools for OSCE station pre-validation, pending replication with multiple runs, larger station banks, and human comparators.

原始摘要(英文原文)· Original abstract
BACKGROUND: Emergency undocking is a high-acuity, low-occurrence event in robotic surgery. Existing robotic curricula do not address this skill, and no validated instrument has been developed to assess performance. We developed the Structured Approach for Fast Emergency Robotic-Undocking (SAFER), a simulation-based assessment tool, and hypothesized it would be reliable, responsive to educational intervention, and able to distinguish between trainee experience levels. We present multisource validity evidence. METHODS: Using a quasi-experimental design, 58 general surgery residents at a single academic institution completed two simulation assessment scenarios-hemorrhage and hemodynamic instability-at two timepoints 6 months apart. Each scenario was evaluated using a 9-item checklist and a global rating scale (maximum score 20 points), scored independently by two trained faculty raters. Between timepoints, participants received a structured teaching session and self-directed learning materials. Validity evidence was gathered across five domains: content, response process, internal structure, relationships to other variables, and consequences of testing. Psychometric analyses included intraclass correlation coefficients (ICC), Cronbach's α, standard error of measurement (SEM), and minimal detectable change at the 95% confidence level (MDC95). RESULTS: Interrater reliability was excellent overall (ICC = 0.95, 95% CI 0.94-0.96) and across both scenarios and timepoints. Internal consistency was acceptable to good (Cronbach's α = 0.72-0.83). The MDC95 was 2.5 points, with 72% of paired learners exceeding this threshold at the second assessment. Mean total score improved by 5.73 points (Cohen's d = 1.38), with largest gains among junior residents. Baseline scores increased significantly with postgraduate year (PGY) level [F(4,102) = 3.99, p = 0.005], supporting known-groups validity. CONCLUSIONS: SAFER, a novel assessment tool for robotic emergency undocking, demonstrated strong reliability, acceptable internal consistency, large responsiveness to training, and expected performance patterns across trainee levels. Further multicenter validation and transfer studies are needed before use in higher-stakes assessments, such as credentialing or advancement decisions.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

SAFER: a simulation-based assessment tool for emergency robotic undocking with multisource validity evidence. — 科研速览 Science Skim