Gündüz Ercan Kutluay, Ahmet Furkan Kumru
Reporting in this literature is moderate to high overall, yet two weaknesses persist: external validation and clinical integration. Work in this area would benefit from closer attention to generalizability, clinical relevance and the existing reporting guidelines.
OBJECTIVE: We set out to examine how completely artificial intelligence (AI)-based orthopedic studies report their methods.
METHODS: A PubMed search covering 2023-2025 was run with a predefined strategy. Screening of titles, abstracts and full texts left 280 eligible papers, and 200 of these were drawn at random for scoring. Each paper was rated against a six-item framework built from the TRIPOD-AI, STARD-AI and CONSORT-AI guidelines.
RESULTS: Mean score across the 200 studies was 4.8±0.9. Twenty-eight percent reached 6/6, 37% scored 5/6 and 23% scored 4/6; the remaining 12% scored 3 or below. Model description, dataset characteristics and performance metrics appeared in every paper (100%), and training and validation procedures in 88%. External validation was far less common, reported by only 32%, while 61% included some form of clinical comparison.
CONCLUSIONS: Reporting in this literature is moderate to high overall, yet two weaknesses persist: external validation and clinical integration. Work in this area would benefit from closer attention to generalizability, clinical relevance and the existing reporting guidelines.