J. Li, H. Rost
Mass spectrometry has become a central technology for lipidomics, with data-independent acquisition (DIA) enabling broad and reproducible sampling of lipid signals. However, the multiplexed fragment-ion spectra in DIA data complicate lipid identification. Here, we introduce OpenLipid, a large language model (LLM)-based workflow for targeted DIA lipidomics. Using assay libraries built from data-dependent acquisition (DDA) results, OpenLipid directly evaluates extracted ion chromatograms (XICs) from DIA data in a zero-shot setting to identify target lipid peaks and generate human-readable rationales for individual lipid identification decisions. We benchmarked OpenLipid against manual annotations across four datasets comprising human plasma and mouse feces analyzed in positive and negative ionization modes. The plasma assay libraries contained 199 target lipids in positive mode and 147 in negative mode. The fecal assay libraries contained 181 target lipids in positive mode and 264 in negative mode. At a 5% false discovery rate (FDR) threshold, OpenLipid identified 110 (55.3%) and 28 (19.0%) library targets in plasma and 130 (71.8%) and 84 (31.8%) in feces, in positive and negative ionization modes, respectively. OpenLipid achieved an overall identification rate comparable to that of DIAMetAlyzer (57.8%, 19.7%, 89.0%, and 17.8% across the corresponding datasets) and substantially higher than that of untargeted MS-DIAL DIA analysis (12.6%, 0.0%, 47.5%, and 0.0%) on the same assay-library targets. LLM-derived chromatographic features also enabled supervised discrimination between correct and incorrect candidate peak groups for target lipids, with median cross-validation average precision values of 0.838-0.912. Together, these results demonstrate that OpenLipid is an effective LLM-based workflow for FDR-controlled targeted analysis of DIA lipidomics data.