Amir Kedan, Henrik Zauber, Meng-Ran Wang, Suyeon Kim, Qionghua Zhu, Liang Fang, Kathryn S Lilley, Wei Chen, Matthias Selbach
Alternative splicing and proteolytic processing expand proteome diversity by generating distinct protein isoforms from a single gene. However, the relationship between transcript isoforms and protein products remains poorly understood because of limitations in current proteomic workflows. Here, we combined full-length mRNA sequencing with protein fractionation and quantitative mass spectrometry to generate an integrated landscape of mRNA and protein isoforms in human RPE-1 cells. To overcome the ambiguity of bottom-up proteomics, we developed IsoFrac, a computational pipeline that resolves protein isoforms from molecular-weight-resolved peptide migration profiles. Using this approach, we identified ∼45,000 full-length transcripts, ∼32,000 open reading frames (ORFs), and ∼14,000 protein isoform candidates. Comparative analyses revealed widespread translation of alternative transcripts and identified shorter protein variants, likely arising from proteolytic processing and/or alternative translation, as a major and underappreciated source of proteome complexity. Our results establish a scalable framework for isoform-resolved proteogenomics and provide a resource for studying protein isoform diversity.