科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Frontiers in Medicine2026-05-08· Consistency (knowledge bases)

Low-energy small language models with retrieval-augmented generation can surpass large-model performance in rheumatology

Sabine Felde, Rüdiger Buchkremer, Gamal Chehab, Christian Thielscher, Jörg H. W. Distler, Matthias Schneider, Jutta G. Richter

原始摘要(英文原文)· Original abstract
Background: Large language models (LLMs) are increasingly explored for clinical decision support but are limited by high computational and energy demands. Smaller language models (SLMs), particularly when combined with retrieval-augmented generation (RAG), may offer a more sustainable alternative. Rheumatology, characterized by diagnostic complexity and guideline-driven management, represents a suitable test domain. Methods: Five state-of-the-art language models (GPT-4o, Mixtral-8 × 7b-32768, Llama-3.1-Nemotron-70b-Instruct, Qwen-Turbo 2.5, Claude-3.5-Sonnet) were evaluated regarding their suitability for clinical decision support using ten standardized, anonymized rheumatology cases. Models were assessed with and without RAG, and with or without a predefined diagnosis. Diagnostic and therapeutic accuracy were quantified using F1 scores. Factual consistency and relevance were assessed using the Retrieval-Augmented Generation Assessment Score (RAGAS). Results: Mixtral-8 × 7b-32768 with RAG achieved the highest diagnostic (72%) and therapeutic (73%) F1 scores. Nemotron-70b showed strong diagnostic performance without RAG (71%), while Qwen-Turbo performed well in therapeutic recommendations without retrieval (72%). The highest RAGAS score was observed for Mixtral with RAG (81%). Performance regarding clinical decision support varied substantially across models and configurations. Conclusion: SLMs combined with RAG can match or exceed the performance of larger LLMs for clinical decision support while requiring significantly fewer computational resources. Despite promising results, clinically relevant errors persisted across all models, underscoring the need for expert oversight and further real-world validation.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Low-energy small language models with retrieval-augmented generation can surpass large-model performance in rheumatology — 科研速览 Science Skim