科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Insects2026-09-04

Evaluating Sequence Foundation Models for Insect Olfactory Gene Discovery: Failure Modes, Evidence Standards, and a Benchmark Agenda.

Huiqin Li, Dian Zhou, Wei Pu, Guoxing Wu, Chun Xiao, Xi Gao, Junfu Yu

原始摘要(英文原文)· Original abstract
Genome-scale searches for insect olfactory genes are vulnerable to rapid receptor divergence, tandem duplication, fragmented gene models, circular labels, and evolutionary leakage between training and test data. Sequence foundation models trained on DNA or proteins may extend homology-based annotation, but available evidence does not support their use as autonomous annotators or direct predictors of sensory function. This critical review organizes the problem around annotation failures rather than model catalogues. We distinguish six inference levels, from candidate locus detection to organism-level function, and define task-specific high-confidence reference criteria. We also propose a benchmark using biological hard negatives, sequence-cluster and taxonomic holdouts, conventional baselines, calibration, abstention, and workload-aware retrieval metrics. Evidence integration is defined operationally as a pipeline that retains model scores alongside similarity, domains, topology, phylogeny, expression, and assay evidence, without promoting a weak signal to a stronger biological claim. A published ESM-2-supported analysis of remote insect chemoreceptor homology illustrates both the potential and the current evidence boundary. A convincing benefit will require an incremental gain over transparent baselines under frozen biologically realistic splits, with reliable uncertainty and a clear route from ambiguous predictions to review or experiment.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Evaluating Sequence Foundation Models for Insect Olfactory Gene Discovery: Failure Modes, Evidence Standards, and a Benchmark Agenda. — 科研速览 Science Skim