Narayanan Raghupathy
Developed ProteoEM, an expectation-maximization framework for weighted proteoform quantification, using fixed, pre-calibrated probe-response rates separate from abundance estimation. In simulations, ProteoEM recovered the true molecular composition, whereas approaches that reduced each affinity trace to a hard yes/no call introduced substantial errors. Grouping traces that carry the same evidence into equivalence classes reduces the EM's work sevenfold without changing the estimates, and ProteoEM corrects for observation yields to estimate the composition of the source sample.
Single-molecule affinity mapping enables measurement of individual proteins and proteoforms, but imperfect and nonspecific probe binding means that affinity traces may be compatible with multiple proteoforms. Accurate abundance estimation therefore requires probabilistically weighting ambiguous traces rather than assigning each trace to a single candidate. We developed ProteoEM, an expectation-maximization framework for weighted proteoform quantification, inspired by transcript abundance estimation methods for RNA-seq and released as an open-source Python package. ProteoEM evaluates each observed affinity trace against all candidate proteoforms using fixed, pre-calibrated probe-response rates that are separate from abundance estimation. It estimates the abundance of each proteoform and reports proteoforms that the probes cannot tell apart as a single group. Because some proteoforms are observed more readily than others, ProteoEM also corrects for these observation yields to estimate the composition of the source sample. We show that grouping traces that carry the same evidence into equivalence classes reduces the EM's work sevenfold without changing the estimates. In simulations, ProteoEM recovered the true molecular composition, whereas approaches that reduced each affinity trace to a hard yes/no call introduced substantial errors. Performance was robust to moderate, uniform calibration error but was biased by informative missing data and by proteoforms absent from the reference set. When observation yields were known, ProteoEM also recovered source-sample composition from observed molecular counts. ProteoEM is an open-source, reproducible framework for quantitative analysis of single-molecule affinity measurements and provides a basis for validation using experimental molecule-level data.