Kees Mandemakers
Linkage methods for constructing historical longitudinal datasets range from manual matching to fully automated procedures, with or without the use of machine-learning techniques or other forms of artificial intelligence. This article examines the extent to which these record-linkage techniques are transparent, repeatable, and consistent, focusing on two aspects that may be regarded as black boxes: the involvement of experts when algorithms fail to produce an unambiguous result, and the use of machine-learning methods based on trained data. At which stages are linkage outcomes shaped by human judgment (beyond the initial formulation of rules and techniques), and what are the strengths and weaknesses of such interventions? The wide variety – and often remarkable ingenuity – of linkage procedures employed across different databases provides an important starting point for this analysis. At the same time, it becomes clear that many publications offer limited transparency when describing specific steps or decision-making processes. Although this critique of black-box elements is fundamental, it does not detract from the substantial effort and impressive creativity with which researchers have addressed the challenge of record linkage. The article concludes with a discussion offering suggestions to establish a stronger foundation for (semi-)manually linked databases and to improve the manual component of the record-linkage process in a more systematic manner. Owing to the limited scope of this article, the analysis focuses on the record-linkage practices of several major databases containing historical nominal data.