科研速览继续刷下去 →
◆ Advanced Computing2026-09-26· Ambiguity

Evaluating Large‐Language Models in Bioinformatics Applications

Hengchuang Yin, Ziwen Cui, Dongxu Li, Yue Yang, Ying Chang, Xiaobo Zhu, Jun Zhang, Pengwei Hu, Xi Zhou, Lun Hu

一句话结论

Large language models (LLMs) demonstrate competitive performance across various bioinformatics tasks with minimal task-specific adaptation. LLMs show limitations in chemical structure reasoning, gene/protein recognition, and computational biology question answering. Integrating LLMs with structured reasoning frameworks could advance next-generation bioinformatics applications.

原始摘要(原文)
ABSTRACT Large language models (LLMs) have significantly revolutionized natural language processing through their strong capabilities in text generation and reasoning. Yet, their applicability to bioinformatics applications remains largely unexplored. Here, we systematically evaluate state‐of‐the‐art LLMs across six representative task domains: drug–drug interaction prediction, antimicrobial and anticancer peptide identification, molecular optimization, gene and protein named entity recognition, single‐cell type annotation, and bioinformatics question answering. Our results show that general‐purpose LLMs can deliver competitive performance across most tasks, demonstrating their versatility for biological data analysis under minimal task‐specific adaptation. Meanwhile, our findings underscore critical limitations for the usage of LLMs in bioinformatics research. First, LLMs exhibit limited capability in chemical structure reasoning for molecular optimization, highlighting the need for integrating physics‐based structure constraints into generative modeling. Second, their sensitivity to ambiguity in gene/protein recognition and single‐cell annotation emphasizes the importance of domain‐specific knowledge in prompt design. Finally, their difficulty in answering computational biology questions suggests that future systems may benefit from combining LLMs with structured reasoning frameworks to support more biologically rigorous inference and decision‐making. These findings reveal both the promise and the boundaries of current LLMs in bioinformatics, offering a roadmap for advancing next‐generation LLM models tailored to the bioinformatics applications.
读原文 ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文

Evaluating Large‐Language Models in Bioinformatics Applications — 科研速览 Science Skim