科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Frontiers in artificial intelligence2026-01-01

Can large language models serve as consultants for forensic cause of death analysis? A multidimensional evaluation.

Enhao Fu, Haojie Qin, Zhiling Tian, Hewen Dong, Donghua Zou, Xiaotian Yu, Ningguo Liu

一句话结论 · In one sentence

DeepSeek-R1 demonstrated a statistically significant advantage in inference quality scores over GPT-4o (p = 0.015, r rb = 0.28) and Gemini-2.5pro (p = 0.000003, r rb = 0.46), while no statistically significant differences were observed among the four models in terms of conclusion accuracy scores. The locally deployed DeepSeek-R1:32b model also showed no statistically significant difference from GPT-4o in conclusion accuracy scores. However, hallucinations persistently appear in the response reports of all LLMs.

原始摘要(英文原文)· Original abstract
INTRODUCTION: Large language models (LLMs) have been proposed as decision support tools in medicine, yet their role in forensic cause of death analysis remains unexplored. METHODS: In this study, we used 118 real-world cases spanning diverse categories of death to systematically evaluate the performance of four representative LLMs (GPT-4o, OpenAI o3, Gemini-2.5pro, and DeepSeek-R1) in forensic cause of death analysis. Two senior forensic pathologists independently evaluated each model's decision-making capabilities regarding inference quality and conclusion accuracy. These metrics were assessed using an expert scoring system with a 5-point Likert scale, with original analytical statements and legally valid expert opinions serving as objective gold standards. In a sub-study, we examined the application potential of the locally deployed open-source model DeepSeek-R1:32b. Additionally, a targeted retrospective analysis was conducted to quantify the incidence and typologies of AI hallucinations. RESULTS: DeepSeek-R1 demonstrated a statistically significant advantage in inference quality scores over GPT-4o (p = 0.015, r rb = 0.28) and Gemini-2.5pro (p = 0.000003, r rb = 0.46), while no statistically significant differences were observed among the four models in terms of conclusion accuracy scores. The locally deployed DeepSeek-R1:32b model also showed no statistically significant difference from GPT-4o in conclusion accuracy scores. However, hallucinations persistently appear in the response reports of all LLMs. DISCUSSION: LLMs can provide limited auxiliary value in cause of death analysis but should not replace the final judgment of forensic experts. LLMs still require expert oversight to ensure evidence integrity and mitigate risks such as hallucination. Open source LLMs can further mitigate data privacy concerns and provide practical support for cause of death analysis.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Can large language models serve as consultants for forensic cause of death analysis? A multidimensional evaluation. — 科研速览 Science Skim