M Frank Erasmus, Daniel Bedinger, Elizabeth Hopkins, Ginger Ferguson, Justine Strickler, Christilyn P Graff, Samantha R Summers, Stacy L Capehart, Joshua D Slocum, Crystal Richardson, Sumit Kumar, Zhifei Sun, Yujie Shang, Jixian Zhang, Ming Gu, Lixia Yi, Alon Wellner, Shuangjia Zheng, Wei Lu, Pietro Sormanni, Matthew Greenig, Haiping Zhang, Brendan T Mann, Mahdi Baghbanzadeh, Ali Rahnavard, Gregory L Moore, Huaiyu Sun, Ying Ding, Alex Nisthal, Jitendra Kanodia, Matthew J Bernett, Aurélien Pélissier, Yanjun Shao, Maria Rodriguez Martinez, Karthik Ramesh, Horacio Nastri, Andreas Evers, Anhar Abdelatif, Andrew J Bordner, Mykola Bordyuh, Lim Heo, Brian A Kidd, H Serhat Tetikol, Shuai Wei, Jung-Eun Shin, Ryan Peckner, Leigh Manley, Ajitesh Lunge, Yashas Devasurmutt, Bora Guloglu, Liviu Copoiu, Miles McGibbon, Monica L Fernandez-Quintero, Nitesh Mishra, Sean M Callaghan, Olivia M Swanson, Daniel L V Bader, James A Ferguson, Sai S R Raghavan, Benjamin Nemoz, Colleen A Maillie, Charles Bowman, Bryan Briney, Andrew B Ward, Paolo Marcatili, Rahmad Akbar, Bing He, Fandi Wu, Jianhua Yao, Bin Hu, Michal Kucer, Kaetlyn Rose Gibson, Rahul Somasundaram, Li-Wei Hung, Tomasz Kaszuba, Daved H Fremont, Hyeongsun Jeong, Vinodh Babu Kurella, Shipra Malhotra, Satyendra Kumar, Yanyun Liu, Lingling Xu, Joshua Misa, Alexander Nicholas St John, Jeff Vogt, Fátima A Dávila-Hernández, Da Xu, Michael Chungyoun, Zyaja D Huggan, Jeffrey J Gray, Jonathan Parkinson, Young Su Ko, Wei Wang, Franziska Geiger, Jonathon D Ziegler, Nikhil Haas, Chance Challacombe, Ahmad Qamar, Akshita Singh, Yi-Ching Tang, Zhiqiang An, Xiaoqian Jiang, Yejin Kim, Xinyan Zhao, Erik Swanson, Jürgen Klattig, Karsten Winkler, Tschimegma Bataa, Volker Sandig, Lilian Denzler, Chunan Liu, Randall J Brezski, Laura Spector, Katheryn Perea-Schmittle, Sara D'Angelo, Fortunato Ferrara, Andrew R M Bradbury
Experimentally validated prospective, blinded benchmarks are needed to separate durable advances from hype in computational antibody design. Here AIntibody, a challenge inspired by the Critical Assessment of Structure Prediction, tests 511 artificial intelligence (AI)-designed or predicted antibodies from 29 organizations on three tasks: in silico affinity maturation from phase 1 sequencing outputs, affinity ranking within heavy-chain complementarity-determining region 3 (HCDR3) clusters of a selection output and CDR design of proteins not included in a selection output. Validated with diverse experimental assays, several groups produced developable antibodies with affinities <100 pM. However, these successes were exceptions that did not transfer across tasks. Affinity-matured antibodies were modeled effectively. Except for one model, predicting high-affinity clones from clustered HCDR3 datasets was worse than random clone picking. Out-of-library design was highly variable for most method submissions, with many failing to outperform standard selections. The AIntibody challenge shows that AI can optimize antibodies in defined, biologically grounded regimes, in addition to highlighting critical gaps including affinity prediction and library-inspired antibody design and cross-task generalization.