Wenhan Wu, Jun Zhang, Zhen Huang
Data-driven lipid nanoparticle (LNP) design increasingly relies on heterogeneous datasets assembled across experimental programs, but random-row validation may overstate performance for unseen sources. We developed a source-aware benchmark using 18,326 functional records, 50 coherent assay strata, and 39 publication-defined source domains. LightGBM, logistic regression, and ridge models were evaluated under random-row, exact-ionizable-lipid, descriptor-cluster, and complete source-domain holdout, with preprocessing confined to training folds and metrics aggregated by source. LightGBM source-macro ROC AUC decreased from 0.764 (95% CI, 0.732-0.796) under random-row validation to 0.585 (0.542-0.628) under source holdout, a paired mean loss of 0.179 (P = 2.5 × 10⁻¹¹). Continuous-rank correlation declined from 0.463 to 0.160, while top-10% candidate enrichment decreased from 2.61-fold to 1.39-fold. Exact-lipid and descriptor-cluster holdouts produced intermediate performance, indicating that chemical novelty did not reproduce complete source novelty. Results were robust to weighting, fold assignment, cluster count, and leave-one-source-out validation. Random-row performance alone therefore did not establish cross-source transportability. Fixed source-aware splits, source-level estimands, and machine-readable outputs provide a reusable framework for more reliable LNP prediction.