Huiqin Li, Dian Zhou, Wei Pu, Guoxing Wu, Chun Xiao, Xi Gao, Junfu Yu
Genome-scale searches for insect olfactory genes are vulnerable to rapid receptor divergence, tandem duplication, fragmented gene models, circular labels, and evolutionary leakage between training and test data. Sequence foundation models trained on DNA or proteins may extend homology-based annotation, but available evidence does not support their use as autonomous annotators or direct predictors of sensory function. This critical review organizes the problem around annotation failures rather than model catalogues. We distinguish six inference levels, from candidate locus detection to organism-level function, and define task-specific high-confidence reference criteria. We also propose a benchmark using biological hard negatives, sequence-cluster and taxonomic holdouts, conventional baselines, calibration, abstention, and workload-aware retrieval metrics. Evidence integration is defined operationally as a pipeline that retains model scores alongside similarity, domains, topology, phylogeny, expression, and assay evidence, without promoting a weak signal to a stronger biological claim. A published ESM-2-supported analysis of remote insect chemoreceptor homology illustrates both the potential and the current evidence boundary. A convincing benefit will require an incremental gain over transparent baselines under frozen biologically realistic splits, with reliable uncertainty and a clear route from ambiguous predictions to review or experiment.