科研速览 · Science Skim继续刷下去 · Keep skimming →
2026-08-01· Judgement

A guide to evaluating LLM-based extraction and judgement tools

Jamie Cummins

原始摘要(英文原文)· Original abstract
Researchers have begun to delegate extraction and judgement research tasks to tools built on large language models (LLMs), including the annotation of text, screening of papers, extraction of numerical values, and identification of supporting passages. Existing guidance establishes that these tools should be validated, documented, and tested for robustness. Researchers still need practical guidance, however, on how precisely different tools can be most appropriately evaluated. This paper provides a step-by-step workflow for evaluating such tools. This involves firstly defining the tool’s output, identifying the desired comparator, and specifying what would count as satisfactory performance – from these decisions, the appropriate analyses and requirements for comparison can be determined and conducted. I provide a worked example of this evaluation approach based on a simulated systematic review extraction-and-judgement tool.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

A guide to evaluating LLM-based extraction and judgement tools — 科研速览 Science Skim