科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Systems2025-10-11· Sustainability

The AI Annotator: Large Language Models’ Potential in Scoring Sustainability Reports

Yue Wu, Peng Hu, Derek Wang

原始摘要(英文原文)· Original abstract
To explore the potential of Large Language Models (LLMs) as AI Annotators in the domain of sustainability reporting, this study establishes a systematic evaluation methodology. We use the specific case of European football clubs, quantifying their sustainability reports based on the sport Positive matrix as a benchmark to compare the performance of three state-of-the-art models (i.e., GPT-4o, Qwen-2-72b-instruct, and Llama-3-70b-instruct) against human expert scores. The evaluation is benchmarked on dimensions including accuracy, mean absolute error (MAE), and hallucination rates. The results indicate that GPT-4o is the top performer, yet its average accuracy of approximately 56% shows it cannot fully replace human experts at present. The study also reveals significant issues with overconfidence and factual hallucinations in models like Qwen-2-72b-instructon. Critically, we find that by implementing further data processing, specifically a Chain-of-Verification (CoVe) self-correction method, GPT-4o’s initial hallucination rate is successfully reduced from 16% to 10%, while accuracy improved to 58%. In conclusion, while LLMs demonstrate immense potential to streamline and democratize sustainability ratings, inherent risks like hallucinations remain a primary obstacle. Adopting verification strategies such as CoVe is a crucial pathway to enhancing model reliability and advancing their effective application in this field.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

The AI Annotator: Large Language Models’ Potential in Scoring Sustainability Reports — 科研速览 Science Skim