S. Dhup, A. Singh
Breast cancer genomics remains disproportionately shaped by cohorts of European and East Asian ancestry, leaving South Asian populations- nearly one-fifth of the world's people- underrepresented despite a distinct clinical presentation: earlier age at onset, later stage at diagnosis, and poorer stage-specific survival relative to Western cohorts. Here we present an integrated multi-omic resource generated from the same 125 patient Indian breast cancer cohort, comprising whole-genome somatic variant calls, paired tumor normal RNA sequencing batches, and quantitative shotgun proteomics, together with clinical annotation spanning age range, menopausal status, tumor laterality, family history, receptor (ER/PR/HER2) status, and recurrence/progression outcome. Somatic profiling recovered canonical breast cancer drivers (TP53, 36%; PIK3CA, 22%; GATA3, 8%; ARID1A, 5%) alongside a set of recurrently altered genes dominated by exceptionally large genes (MUC4, TTN, OBSCN, USH2A), a recognized signature of length-driven passenger-mutation recurrence. Cross-referencing recurrent protein-level alterations against the OncoKB knowledge base identified PIK3CA/AKT1 pathway hotspot mutations, most frequently PIK3CA H1047R, in one-fifth of patients (25/125). Transcriptomic analysis cleanly separated tumors from normal tissue along the first principal component and supported a curated immune gene panel capable of stratifying tumors by immune-associated expression pattern. We provide this dataset as a resource for driver and actionability benchmarking, cross-platform (DNA-RNA-Protein) concordance analysis, tumor-immune stratification, and methodological work on length-normalized driver detection in an Indian-ancestry context that remains largely absent from reference cancer genomics datasets.