WANG Pengyu, JIANG Shuming, WEI Zhiqiang, YU Jun, ZHANG Mengmeng
Current statistical work in the field of educational development faces challenges such as low knowledge-retrieval efficiency and high technical barriers to data utilization. This study focuses on the knowledge related to the indicator system that underpins educational statistical work and proposes a large language model (LLM)-based method for constructing a knowledge graph for the educational statistical indicator system so as to provide knowledge support for addressing the above issues. Specifically, a research framework comprising the following layers is established: data processing, ontology construction, graph construction, and application prospect. In the data processing layer, original documents are converted into Markdown format and cleaned to enhance structural parsing. In the ontology construction layer, a seven-step method is used to construct the ontology of the educational statistical indicator system. In the graph construction layer, an enhanced prompting strategy integrating chain-of-thought prompting with the ontology structure guides the LLM in knowledge extraction, with Neo4j used for storage and visualization. Experimental results show that this strategy achieved an F1 score of 96.22% for entity-attribute extraction and 92.23% for entity-relation extraction, significantly outperforming the basic prompting strategy, thereby verifying the effectiveness of chain-of-thought prompting in complex information extraction. The constructed knowledge graph effectively represents the complex knowledge in the educational statistical indicator system, providing a knowledge foundation for intelligent information retrieval and natural language data querying and offering insights for knowledge organization in other statistical fields.