Weijian Chen, Haoxiang Xu, Daojian Cheng
The vast majority of experimental knowledge in chemistry and materials science remains locked in unstructured natural language texts, creating a fundamental bottleneck for data-driven discovery.Large language models (LLMs), with their remarkable in-context learning and instruction-following capabilities, offer a transformative solution by enabling rapid, scalable conversion of scientific literature into structured databases.This perspective critically examines the end-to-end workflow of LLM-based data extraction, from preprocessing and LLM interaction to postprocessing and validation, while highlighting the unique advantages of chemical domain knowledge for constraining and verifying outputs.We systematically identify hidden risks in large-scale extraction, including condition-value mismatches, error accumulation, "ingratiation" bias and instruction ambiguity, drawing on practical observations.To move toward reliable and robust extraction, we propose four engineering principles, including task decomposition, multilayer validation, clear division of labor between LLMs and deterministic tools, and standardized data reporting in publications.Finally, we outline future research frontiers, namely multimodal agents, crossdocument integration and lowresource benchmarks, that will shape the next generation of autonomous scientific data assistants.