科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ iScience2026-09-18

Systematic assessment of text summarization methods for biomedical literature from frequency methods to language models.

Fabio Baumgärtel, Enrico Bono, Lucas Fillinger, Louiza Galou, Kinga Kęska-Izworska, Samuel M Walter, Peter Andorfer, Klaus Kratochwill, Paul Perco, Matthias Ley

原始摘要(英文原文)· Original abstract
The rapid expansion of biomedical literature demands automated summarization tools that reliably condense research articles into concise, accurate summaries. We benchmarked 62 summarization methods, ranging from frequency-based and TextRank extractors to encoder-decoder models (EDMs) and large language models (LLMs), on 1,000 biomedical abstracts from 20 journals across ScienceDirect and Cell Press, using author-written highlights as reference summaries. Models were evaluated with a composite suite of lexical, semantic, and factual metrics, including ROUGE, BLEU, METEOR, embedding-based similarity, and factuality scores. General-purpose models (e.g., Mistral, GPT, and Llama) achieved the highest overall performance across lexical and semantic dimensions, outperforming reasoning-oriented (e.g., DeepSeek and Magistral) and domain-specific (e.g., BioGPT and BioMistral) models. Notably, medium-sized models outperformed large-scale models, suggesting an optimal balance between model capacity and efficiency, while classical extractive methods lagged behind neural approaches. These findings provide a systematic reference for selecting biomedical summarization tools and highlight that broad pretraining outperforms narrow domain adaptation.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Systematic assessment of text summarization methods for biomedical literature from frequency methods to language models. — 科研速览 Science Skim