科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ IEEE transactions on computational biology and bioinformatics2026-09-25

Protein FT-Transformer: Species Classification via Amino Acid-Encoded DNA Sequences.

Meetesh Nevendra, Deepak Singh

一句话结论 · In one sentence

Each DNA barcode is translated codon by codon into its corresponding amino acid sequence, preserving protein-level constraints that remain informative for species discrimination, providing a compact and interpretable input for learning. The proposed Protein FT-Transformer, an amino-acid-based transformer architecture for species classification, employs a dual-stream feature tokenizer and a multi-head self-attention encoder to model dependencies across distant sequence positions. The proposed framework is evaluated on six publicly available DNA barcode benchmarks spanning invertebrate, vertebrate, and plant taxa and outperforms classical machine learning methods.

原始摘要(英文原文)· Original abstract
DNA barcoding is widely used for species identification, yet many computational classifiers continue to operate directly on raw nucleotide sequences. For large and heterogeneous barcode repositories, this representation can introduce avoidable computational cost because long nucleotide strings often contain redundant codon-level information and highly uneven class distributions. This study investigates whether a biologically grounded representation can improve supervised barcode classification while reducing the effective input length. Each DNA barcode is translated codon by codon into its corresponding amino acid sequence, preserving protein-level constraints that remain informative for species discrimination. The translated sequence provides a compact and interpretable input for learning, without requiring large-scale pretraining or an additional decompression stage. On this representation, we propose Protein FT-Transformer, an amino-acid-based transformer architecture for species classification. The model employs a dual-stream feature tokenizer to combine residue-level information with broader contextual features, followed by a multi-head self-attention encoder that models dependencies across distant sequence positions. Training further incorporates class-weighted focal loss to address class imbalance, class-wise temperature scaling to improve probability calibration, and lightweight ensemble classification heads to improve prediction stability. The proposed framework is evaluated on six publicly available DNA barcode benchmarks spanning invertebrate, vertebrate, and plant taxa. Protein FT-Transformer outperforms classical machine learning classifiers and nucleotide-based deep learning baselines, achieving an average classification accuracy of 98.20% with improved F1-score and AUC values under modest computational and memory requirements. t-SNE visualizations of the learned embeddings also show more coherent species-level grouping. The results suggest that codon translation coupled with a task-specific transformer provides an effective supervised strategy for DNA barcode-based species identification.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Protein FT-Transformer: Species Classification via Amino Acid-Encoded DNA Sequences. — 科研速览 Science Skim