科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Research synthesis methods2026-09-21

Workflow-induced evaluation targets in automation-assisted screening: Implications for cumulative evidence in research synthesis.

John Alan Nunnery

原始摘要(英文原文)· Original abstract
Automation-assisted title-abstract screening is routinely evaluated using metrics such as recall, precision, and workload reduction, and results are increasingly summarized across studies. Synthesis presupposes that evaluation results refer to comparable target quantities. This article examines whether screening evaluation results produced under diverse automation-assisted workflows can meaningfully support cumulative inference, and under what conditions cross-study comparison is warranted. This conceptual article treats screening evaluation targets as workflow-induced rather than task-inherent. Evaluation design is decomposed into four workflow dimensions (representation, inference, governance, and evaluation) plus the reference-decision set produced or specified. Recurring choice patterns across these elements define evaluation configurations. Two study-level configurations and one synthesis-level pattern are illustrated: endogenous adjudication, inherited adjudication, and aggregation of heterogeneous targets. Across configurations, evaluation targets are induced by workflow-specific reference-decision sets and, in some designs, model-based target domains rather than by a common reference set. Nominally similar metrics can therefore estimate different quantities across studies, and apparent cumulativeness can arise from pooling estimates of different quantities. Standard heterogeneity statistics cannot distinguish variation in method effectiveness from variation in what is being estimated. Comparison is warranted only when targets align before pooling. Differences across screening evaluations may reflect differences in what is being estimated rather than statistical heterogeneity around a common estimand. When evaluation targets are not aligned across the framework, aggregate summaries can describe reported results but cannot support cumulative inference. Improved aggregation requires both better study-level specification of evaluation targets and explicit attention to configuration variation at synthesis .
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Workflow-induced evaluation targets in automation-assisted screening: Implications for cumulative evidence in research synthesis. — 科研速览 Science Skim