科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ ACS Catalysis2025-10-20· Computer science

Distilling Knowledge from Catalysis Literature with Long-Context Large Language Model Agents

Honghao Chen, Hongxuan Liu, Yishen Tew, Xiaotian Ren, Xiaojin Tang, Xiaonan Wang

原始摘要(英文原文)· Original abstract
Decades of catalysis knowledge remain locked in unstructured prose, hindering data-driven discovery. Existing text-mining tools struggle to establish the synthesis–structure–performance relationships critical for catalyst knowledge discovery as they rarely connect synthesis protocols in one section with the resulting material properties and performance outcomes reported elsewhere. Here, we present CATDA (Corpus-aware Automated Text-to-Graph Catalyst Discovery Agent), a long-context large language model-driven agentic framework that reads full documents and distills them into actionable, provenance-tracked knowledge graphs linking material properties, multistep synthesis, conditions, and testing outcomes. Applied at corpus scale, CATDA extracts data with near-human fidelity (F1 = 0.983) and a 12-fold speedup over manual curation. This structured knowledge is made accessible through two synergistic applications: a DatasetAgent for exporting machine-learning-ready tables and a CatAgent providing a conversational, citation-linked interface for interactive discovery. The high-quality data set enabled the training of a predictive model for ethylbenzene conversion while simultaneously exposing systemic challenges such as feature sparsity and protocol heterogeneity in the source literature. By transforming the literature into a queryable and computable resource, CATDA offers a scalable route to accelerate large-scale data analysis, quantitative modeling, and a rational catalyst design paradigm.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Distilling Knowledge from Catalysis Literature with Long-Context Large Language Model Agents — 科研速览 Science Skim