科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Computational Materials Science2026-07-31· Bottleneck

Information extraction from literature for ORR catalyst in fuel cell

Hein Htet, Manae Hirano, Amgad Ahmed Ali, Yutaka Sasaki, Ryoji Asahi

原始摘要(英文原文)· Original abstract
The oxygen reduction reaction (ORR) catalyst is critical for fuel cell efficiency, yet extracting structured technical data from the rapidly expanding volume of scientific literature remains a bottleneck for material discovery. In this study, we propose a high-performance framework utilizing the DyGIE++ multi-task learning architecture with domain-specific BERT variants to construct the Fuel Cell Corpus for Materials Informatics (FC-CoMIcs). Using a curated dataset of 76 expert-annotated articles (8898 sentences), we fine-tuned multiple models and identified our top-performing MatSciBERT-based model to maximize extraction precision across complex scientific spans. Experimental evaluations on the standard test set yielded an NER F1-score of 74.82% and an RE F1-score of 62.04% using our best-performing model (model-4, fine-tuned on MatSciBERT-3). To rigorously validate this approach, we conducted a fair comparison in which model-4 and five frontier Large Language Models (LLMs) released in early 2026 were evaluated on the same 200-sentence random sample drawn from the test set. Model-4 achieved NER F1 of 76.19% and RE F1 of 66.58%, significantly outperforming the best-performing LLM (Gemini 3.1 Pro Preview) by +19.1% in NER and +29.8% in RE, thereby demonstrating that specialized fine-tuned architectures remain superior to general-purpose reasoning models for technical materials data extraction. To demonstrate industrial-scale utility, the framework was applied to a corpus of 10,000 articles (2010–2024), extracting over 4.3 million entities and 4.2 million relations. By implementing the three-step refinement pipeline (unit filtering, numeric validation, and normalization), we conducted a longitudinal analysis of “power density” trends. This analysis captured the performance evolution of catalyst families, specifically highlighting the transition from traditional Pt-based systems to emerging high-performance non-metal alternatives. This framework transforms unstructured literature into a high-resolution structured database, providing a powerful foundation for accelerated catalyst design.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Information extraction from literature for ORR catalyst in fuel cell — 科研速览 Science Skim