M. Wang, S. Verma, S. Jayasundara, S. D. Kadadi, M. Kazemian, A. Grama, N. A. Lanman
Background: Effectively translating data on complex cellular responses from transcriptomic and morphological measurements into molecular design remains a significant computational challenge. Existing generative methods operate on single modalities and condition on post-treatment measurements without leveraging paired control-treatment dynamics to capture perturbation effects. Results: We present Pert2Mol, a framework for multi-modal phenotype-to-structure generation that integrates transcriptomic and morphological features from paired control-treatment experiments. Pert2Mol employs bidirectional cross-attention between control and treatment states to capture perturbation dynamics, conditioning a rectified flow transformer that generates molecular structures along straight-line trajectories. We introduce Student-Teacher Self-Representation (SERE) learning to stabilize training in high-dimensional multi-modal spaces. On the Ginkgo Data Platform (GDP) dataset, Pert2Mol achieves Frechet ChemNet Distance of 4.996 compared to 7.343 for diffusion baselines and 59.114 for transcriptomics-only methods, while maintaining perfect molecular validity and appropriate physicochemical property distributions. The model demonstrates 84.7% scaffold diversity and 12.4 times faster generation than diffusion approaches with deterministic sampling suitable for hypothesis-driven validation. Conclusions: Pert2Mol establishes a new paradigm for linking high-content phenotypic screening data with computational hypothesis generation in drug discovery: joint multi-modal perturbation modeling enables more accurate and structurally diverse molecular design than single-modality or diffusion-based approaches. Code and pretrained models are available at https://github.com/wangmengbo/Pert2Mol