Dmytro Zaharnytskyi
Biosecurity screening of synthetic DNA relies on sequence-homology tools such as BLAST, motivating interest in learned, reference-database-free classifiers as a complement. We present, to our knowledge, the first benchmark comparing a nucleotide foundation model (DNABERT-2) against a protein foundation model (ESM-2) for virulence-factor classification under a leave-one-pathogen-out holdout, on a single consumer GPU. Across 13 organisms (one designated reference assembly each-except a partial o=S. cerevisiae, and six-strain Ebola; seven pathogens held out in turn; 64,803 in-frame, CDS-derived windows), DNABERT-2,fine-tuned with LoRA fails to generalize to held-out pathogens (mean F1 = 0.18, AUC = 0.54)-no better than a k-mer logistic-regression baseline (AUC = 0.51). ESM-2 fine-tuned on the translated windows performs substantially better (mean F1 = 0.54, AUC = 0.77), exceeding DNABERT-2 on every fold (paired Wilcoxon p=0.016 for F1, p=0.031 for AUC). The ESM-2-DNABERT-2 gap replicates within a second, independently assembled eight-pathogen benchmark after retraining (ESM-2 AUC ≈ 0.92 vs. 0.68), although direct cross-dataset transfer is poor (ESM-2 AUC = 0.59) and the external labels are homology-defined, which may favor a protein encoder. A BLASTx stress test detects ≥99% of mutated pathogenic queries up to 70% conservative substitution, using computationally simulated (not function-validated) substitutions against a small custom database. We read these findings as evidence within this benchmark-not as a general claim that protein representations are superior for biosecurity screening, and not as an operational screen: even ESM-2 is miscalibrated and low-precision under realistic class imbalance. All code runs on one NVIDIA RTX 3090.