科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ The Electronic Library2025-12-19· Metadata

Describing archival photographs using multimodal LLMs: a case study on evaluating vision-language model performance for creating descriptive metadata

Drew Facklam, Sarah Sweeney, Shoumik Majumdar, Rahul Kamath

原始摘要(英文原文)· Original abstract
Purpose This paper aims to examine the efforts of librarians in Northeastern University Library’s Digital Production Services department to employ Gemini, a pre-trained multimodal large language model, for generating descriptive metadata for an archival photographic collection. Design/methodology/approach The project comprised three phases: 1) researching and selecting a multimodal LLM that was both accurate and cost-effective for generating descriptive metadata for photographs; 2) developing an application to batch-submit photographs to the chosen model’s API and convert the results into a human-readable spreadsheet and 3) evaluating the output for completeness, accuracy, consistency and potential bias. Findings Insights from model research guided the development and deployment of an application that queried the Gemini model to generate titles, abstracts and other key metadata. Iterative testing showed that although accuracy, completeness and consistency were imperfect, the output quality demonstrated strong potential for future implementation. Analysis of application results indicates that ongoing errors and biases may be reduced through strategic prompt engineering and systematic quality control measures. These testing methodologies will inform future efforts to operationalize computer vision workflows for processing archival photographs. Originality/value While computer vision holds significant potential to improve cataloguing workflows for archival photograph collections, few case studies have explored methods for evaluating and operationalizing these workflows to produce suitable metadata records. This paper presents an evaluation process to assess completeness, accuracy, consistency and potential bias in the generated metadata. Crucially, this workflow is adaptable and can be repeated as the application, prompts and models evolve, ensuring ongoing reliability and improvement.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Describing archival photographs using multimodal LLMs: a case study on evaluating vision-language model performance for creating descriptive metadata — 科研速览 Science Skim