K Wolcott, J Hammock, K Schulz, T M Orrell
Natural history museums house invaluable biodiversity archives, but discoveries are often limited by the accessibility of collections data. As digitization of specimens and metadata accelerates, automated methods are needed to organize image databases and enable large-scale analyses. Beyond species identification, computer vision (CV) can automate image database management by standardizing gallery displays through intelligent cropping and extracting tags to improve search functionality. Encyclopedia of Life is a biodiversity database that aims to document all ∼2.2 million known species. Using EOL images and crowdsourced data, combined with pretrained and custom-trained models, we present 14 CV pipelines in three categories: Object Detection for Image Cropping, Object Detection for Image Tagging, and Classification for Image Tagging. Pipelines are open source, available on GitHub (https://github.com/EOL/computer-vision-with-EOL-images), and documented for nonexperts. We share models, datasets, and methodological insights demonstrating how CV can unlock novel information from over 1.5 million images.