Valentina Ramundi, Michael Witting
Liquid chromatography-mass spectrometry (LC-MS)-based non-targeted metabolomics produces intricate datasets that need advanced tools for identifying metabolites (MetID). Metabolite annotation and identification in non-targeted metabolomics require high-quality fragmentation data from biological samples and reference libraries. Public and commercial databases, such as NIST, MassBank, MoNA, GNPS, HMDB, mzCloud, and METLIN, are vital resources for both spectral matching and providing training data for machine-learning-driven MetID tools. These libraries differ in their coverage, curation, and access models, and they are often supplemented by in-house databases that cater to specific laboratory conditions, ensuring the highest level of confidence in identifications. To maximize the benefits of MS2 libraries and ensure they work seamlessly together, we rely heavily on standardized file formats. Standard formats, such as mzML, MGF, MSP, JSON, and the MassBank format, each come with different levels of metadata richness and compatibility with various software tools. This chapter gives an overview of key MS2 libraries, discusses the strengths and weaknesses of standard data formats, and introduces R-based solutions for better integration.