Sureka Rajendran, Ambily Sivadas, Neeraj Sidharthan, Georg Gutjar, Kanaga Ganesan, Anupam Joshi, Ankur Chaturvedi, Prakash Sadashiv Gambir, Nagarajan M Phani
Human leukocyte antigen matched unrelated donor (URD) identification remains a major challenge for Indian patients due to high genetic diversity and endogamy. Utilizing a HLA high-resolution dataset of 30,027 umbilical cord blood (UCB) units, this study modeled 10/10 and 8/8 matching probabilities for five broad regions and nine sub-regional linguistic populations. Two statistical frameworks, the Formula Matching Method (FMM) and the Actual Measurement Method (AMM) were used to project match rates for registry sizes ranging from 5000 to 1000,000 donors. Using FMM, the probability of identifying a 10/10 and 8/8 matched donor in sub-regional cohorts was approximately 25% at a registry size of 30,000 and increased substantially with larger registry size, reaching approximately 40% at a registry size of 100,000. By relaxing the matching criteria to 9/10 or 7/8, 50% match probability was achieved with a registry size of 30,000. Conversely, AMM produced more conservative estimates, highlighting the effect of stochastic sampling and registry composition. Importantly, significant regional disparities were observed, with highly diverse populations requiring disproportionately larger registries. These findings are highly relevant for determining the optimal national registry size and planning a geographically targeted recruitment by prioritizing highly diverse regions.