Nyeong-Jin Cheon, Xuan Cuong Nguyen, Tatsuya Unno
Whole-genome sequencing has become the standard molecular platform for antimicrobial resistance surveillance, yet existing computational approaches-gene catalogs that ignore expression, k-mer machine learning (ML) that lacks mechanistic interpretability, and Single Nucleotide Polymorphism-based methods that capture only one resistance mechanism-suffer from systematic genotype-phenotype discordance whose molecular causes remain unquantified. We present a 16-step genomic annotation protocol that integrates promoter strength, ribosome binding site efficiency, codon adaptation, mobile element context, and chromosomal point mutations, and benchmark four ML classifiers across 20 296 bacterial genomes from five World Health Organization priority species with matched susceptibility phenotypes. Variant-specific discordance analysis revealed three categories of genotype-phenotype mismatch: antibiotics predictable by gene presence alone (tetracyclines, sulfonamides); those requiring expression context to distinguish functional from silent genes (β-lactams, aminoglycosides); and those dependent on chromosomal mutations invisible to gene catalogs (quinolones). Silent resistance genes showed 30%-85% lower predicted translation initiation rates, 4%-14% lower gene coverage, and weaker promoter activity than functional copies. Among the four classifiers evaluated, Random Forest achieved the highest performance. Expression features improved its F1 by 0.07-0.11 for β-lactams and aminoglycosides but were irrelevant for quinolones, where point mutations were essential. Mobility and regulation features were dispensable. The full model achieved mean F1 of 0.804-0.932 across species; SHapley Additive exPlanations analysis confirmed that the dominant predictive feature per drug class directly reflects its known resistance mechanism, and phylogeny-blocked cross-validation confirmed that single-gene predictions generalize across lineages. These findings establish a mechanistic framework for when and why genomic prediction fails, with direct implications for the design of clinical sequencing-based diagnostic tools.