科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Diagnostic and Interventional Radiology2026-02-26· Medicine

Reporting checklist for foundation and large language models in medical research (REFINE): an international consensus guideline

Ismail Mese, Tugba Akinci D’Antonoli, Christian Bluethgen, Keno Bressem, Renato Cuocolo, Akshay Chaudhari, Ali S. Tejani, Amanda Isaac, Andrea Ponsiglione, Aymen Meddeb, Bardia Khosravi, Bastien Le Guellec, Charles E. Kahn, Chong Hyun Suh, Daniel Pinto dos Santos, Dow-Mu Koh, Eleftherios Tzanis, Elmar Kotter, Errol Colak, Felipe Kitamura, Felix Busch, Felix Nensa, Guang Yang, Henning Müller, Jakob Nikolas Kather, Jawed Nawabi, Jens Kleesiek, Jingyu Zhong, João Santinha, Johannes Haubold, José Guilherme Almeida, Karim Lekadir, Kostas Marias, Lara Noelle Reiner, Lena Maier- Hein, Linda Moy, Lisa C. Adams, Luis Martí- Bonmatí, Magdalini Paschali, Mana Moassefi, Matthias Dietzel, Merel Huisman, Michael Ingrisc, Michail E. Klontzas, Nikolaos Papanikolaou, Oliver Diaz, Paulo Kuriki, Philipp Seeböck, Pouria Rouzrokh, Quirin D. Strotzer, Seong Chan Park, Shahriar Faghani, Soroosh Tayebi Arasteh, Su Hwan Kim, Vasantha Kumar Venugopal, Woojin Kim, Burak Kocak

原始摘要(英文原文)· Original abstract
PURPOSE: To develop the REporting checklist for FoundatIon and large laNguagE models (REFINE), an international reporting guideline for transparent and reproducible reporting of foundation model (FM) and large language model (LLM) studies in medical research, including imaging artificial intelligence (AI) applications. METHODS: The protocol was prespecified and publicly archived. A modified Delphi process was conducted to establish reporting standards for unimodal and multimodal FM and LLM applications involving text, imaging, and structured data. The steering committee coordinated protocol development, expert recruitment, all Delphi rounds, and the harmonization phase. Decisions were made based on predefined consensus thresholds. In Rounds 1 and 2, structured ratings and free-text feedback informed iterative revisions. In the post-Delphi harmonization phase, terminology was standardized, and detailed reporting instructions were finalized. RESULTS: The REFINE development group comprised 57 contributors from 17 countries, and 54 panelists from 16 countries completed Rounds 1 and 2. The harmonization phase was completed by three expert panelists and the steering committee. The entire process produced a 44-item, six-section framework with standardized terminology and detailed reporting instructions, supported by an online platform for practical use (https://refinechecklist.github.io/refine/checklist.html). CONCLUSION: The REFINE provides a comprehensive, consensus-based reporting standard for medical FM and LLM research, including imaging AI studies. The online version facilitates practical implementation. CLINICAL SIGNIFICANCE: The REFINE enables transparent, comparable, and reproducible reporting of FM and LLM studies, supporting reliable evidence synthesis in medical and imaging-focused AI studies.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related