Y. Qu, X. Liu, X. Wang, M. Chen, X. Luo, L. Lyu, M. Yin, J. Hui, D. Yin, R. Dinesh, L. Qiu, K. Huang, H. Wang, S. Tong, H. Cousins, R. Feng, O. Martinez, J. Zhang, T. Chen, R. Altman, J. Leskovec, A. Regev, M. Wang, L. Cong
Target discovery in functional genomics remains largely manual and time-consuming, lacking systematic tools for efficient and reproducible gene-level hypothesis generation. We introduce AutoScreen, an AI co-scientist system supporting target discovery through both Pre-screen Design, which constructs perturbation libraries de novo from free-text research descriptions, and Post-screen Analysis, which re-ranks experimental screen hits by integrating statistical scores with biological context. AutoScreen leverages the complementary strengths of multiple specialist agents that perform deep research into multi-modal evidence, information restructuring, parallel searches across >26 biomedical databases, evidence synthesis, and target review to provide transparent, explainable gene prioritization with pipeline provenance. Across 320 genome-scale CRISPR screens as expert-curated benchmarks, AutoScreen achieved a ~19% increase in validated-hit recovery among its top 100 predictions, relative to the strongest agent baseline, and required a 1.2-fold smaller library to recover the same number of hits at the top-500 reference point. AutoScreen reached mean average precision more than two orders-of-magnitude above random baseline. Further, we validated AutoScreen in cancer immune-evasion case studies focusing on natural killer (NK) and T-cell therapeutics. AutoScreen recovered NK-resistance genes that were initially lower-ranked in a leukemia screen, moving MUC1, PDPN, and LRRC15 from raw ranks of 118, 81, 1384 to 5, 44, 659, respectively. In follow-up tumor killing assay with primary human NK cells, individual perturbation validated all three hits successfully. Next, in prospective benchmarking across two cytotoxic T-cell-killing screens, AutoScreen recovered 77.1% of ground-truth hits called by the consensus of two gold-standard analysis pipelines (FDR<0.1), achieving 8% improvement over the strongest general-purpose LLM agent baseline. Finally, we re-analyzed >500 public datasets to construct the AutoScreen Resource Hub, a growing knowledge base of pre-computed, annotated reports for CRISPR screens, RNA-seq differential expression, and gene-level UK Biobank genome-wide association studies. This agentic AI approach enables real-time genomics benchmark construction that are continuously updated, expanded, for evaluating frontier AI co-scientists. Overall, AutoScreen enables AI-powered target discovery to be more auditable, scalable, extensible, and reproducible.