David Purkarthofer, Sebastian Labenbacher, Helmar Bornemann-Cimenti, Giovanni Landoni
Systematic reviews and meta-analyses aim to comprehensively identify, summarize and appraise literature. Record identification is performed within multiple databases with important overlap. Deduplication is therefore an important, time-consuming step often underreported or described only briefly. Unique identifiers like the digital object identifier (DOI®) provide a practical basis to identify and remove unambiguous duplicates, but require careful handling of missing, inconsistently formatted, or non-unique DOIs. We created a simple tool that removes duplicates based on normalized DOIs and titles, and make it available as a free-to-use web application hosted on deduplicate.it. It is designed for ease of use, while maintaining high transparency by utilizing a human-readable, simple algorithm and outputting exclusion files and a PRISMA-style flowchart alongside the deduplicated output. In a validation set of five reviews in medicine, education and social welfare, deduplicate.it was able to identify around 80% of duplicates and reduce the time dedicated to deduplication accordingly.