科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Archives of Toxicology2026-05-14· Computer science

Transforming animal study toxicology reports into structured, harmonized data using large language models

Tatyana Y. Doktorova, Ilya Schneider Chernov, Alberto Formaggio, Xiaochen Zheng, Alvaro Serra, Preethi Latha Bhavani Mohan, Guillemette Duchâteau-Nguyen, Dragomir Draganov, Matt Hayes, Catrin Hasselgren, Vishal B. Siramshetty, Stefano Gaudio, Marco Tecilla, Pierre Maliver, Eunice Musvasva, Lennart T. Anger

原始摘要(英文原文)· Original abstract
Preclinical toxicology study reports contain the expert interpretations required to distinguish test article-related effects from incidental findings, yet these conclusions often remain embedded in unstructured text form that limit systematic reuse and integration with computational safety approaches. To address this gap, we developed a Large Language Model (LLM) - supported pipeline that converts toxicology reports into structured, machine-readable datasets harmonized with SEND terminology. The pipeline combines automated document preprocessing, section identification, schema-constrained information extraction, and semantic harmonization, complemented by targeted human curation. We evaluated the system performance using 200 Roche toxicology study reports, encompassing clinical pathology, histopathology, organ weights, exposure data, and study-level conclusions. Across domains, extraction performance was strong, characterized by consistently high sensitivity and precision for most parameters. Histopathology, organ weight, and NOAEL-related endpoints demonstrated the greatest robustness, with sensitivity typically above 95% and precision frequently exceeding 97%. Lower performance for parameters such as route of administration and substance identifiers reflected heterogeneous reporting practices rather than LLM-based method limitations. The structured datasets generated by this pipeline enable cross-study querying, identification of compounds with defined toxicological liabilities, integration with raw SEND data, and development of high-quality labels for predictive toxicology models. Practical utility has been demonstrated through representative use cases. These results demonstrate that LLM-assisted extraction can reliably capture expert toxicological interpretations at scale and provide a foundation for data-centric safety assessment, strategic decisions, reverse and forward translational toxicology research.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Transforming animal study toxicology reports into structured, harmonized data using large language models — 科研速览 Science Skim