Meetesh Nevendra, Deepak Singh
Each DNA barcode is translated codon by codon into its corresponding amino acid sequence, preserving protein-level constraints that remain informative for species discrimination, providing a compact and interpretable input for learning. The proposed Protein FT-Transformer, an amino-acid-based transformer architecture for species classification, employs a dual-stream feature tokenizer and a multi-head self-attention encoder to model dependencies across distant sequence positions. The proposed framework is evaluated on six publicly available DNA barcode benchmarks spanning invertebrate, vertebrate, and plant taxa and outperforms classical machine learning methods.
DNA barcoding is widely used for species identification, yet many computational classifiers continue to operate directly on raw nucleotide sequences. For large and heterogeneous barcode repositories, this representation can introduce avoidable computational cost because long nucleotide strings often contain redundant codon-level information and highly uneven class distributions. This study investigates whether a biologically grounded representation can improve supervised barcode classification while reducing the effective input length. Each DNA barcode is translated codon by codon into its corresponding amino acid sequence, preserving protein-level constraints that remain informative for species discrimination. The translated sequence provides a compact and interpretable input for learning, without requiring large-scale pretraining or an additional decompression stage. On this representation, we propose Protein FT-Transformer, an amino-acid-based transformer architecture for species classification. The model employs a dual-stream feature tokenizer to combine residue-level information with broader contextual features, followed by a multi-head self-attention encoder that models dependencies across distant sequence positions. Training further incorporates class-weighted focal loss to address class imbalance, class-wise temperature scaling to improve probability calibration, and lightweight ensemble classification heads to improve prediction stability. The proposed framework is evaluated on six publicly available DNA barcode benchmarks spanning invertebrate, vertebrate, and plant taxa. Protein FT-Transformer outperforms classical machine learning classifiers and nucleotide-based deep learning baselines, achieving an average classification accuracy of 98.20% with improved F1-score and AUC values under modest computational and memory requirements. t-SNE visualizations of the learned embeddings also show more coherent species-level grouping. The results suggest that codon translation coupled with a task-specific transformer provides an effective supervised strategy for DNA barcode-based species identification.