F. Dai, S. You, Y. Zhu, Y. Gao, L. Fu, X. Zhou, J. Su, C. Wang, Y. Fan, X. Ma, X. Deng, L. Yu, H. Qian, Y. He, Y. Ke, C. Han, X. Chang, L. Zheng, S. Wang, Y. Wang, A. Zeng, S. Wang, T. Si, J. Liu, H. Lu, F. Yuan
Programming biological function-designing bespoke proteins to perform specified molecular tasks-is a foundational goal of molecular engineering. However, current design paradigms remain fundamentally limited: they typically require either natural proteins as starting points for optimization or manual reformulation of functional goals as geometric and sequence-level constraints to guide candidate generation. Here we introduce Pinal, a 16-billion-parameter foundation model that designs candidate proteins from natural-language descriptions of desired function. Trained on 1.7 billion synthetically annotated protein-text pairs, Pinal links functional intent to protein sequence and structure. In computational evaluations, generated candidates combined high predicted foldability with functional-description alignment and sequence diversity, providing a basis for prioritizing experimentally testable designs. We applied Pinal to four distinct design tasks-a fluorescent protein, a polyethylene terephthalate hydrolase, an alcohol dehydrogenase and a metabolic H-protein-and experimentally observed the intended function in each case, including catalytic activity for both designed enzymes. Crucially, without iterative experimental optimization, a Pinal-designed H-protein increased product titer by 1.7-fold relative to the corresponding native E. coli H-protein in a multi-enzyme CO2 fixation pathway. These findings support natural language as a high-level interface for candidate generation in protein design, enabling programmable exploration with reduced reliance on manually specified structural or sequence constraints.