Shweta Rana, Pranshu Bhagat, Debnath Pal, Sanghamitra Pati, Harpreet Singh
AI is rapidly expanding in health care. However, there is a significant underrepresentation of low- and middle-income countries (LMICs) in datasets used to train AI applications. Current Food and Drug Administration (FDA)-approved AI tools predominantly use data from high-income countries, with fewer than 4% reporting geographic or racial diversity. This geographic skew leads to significant performance degradation when these tools are used in LMIC populations, perpetuating an equity crisis where health burdens are highest., Addressing this disparity, we highlight the Medical Imaging Datasets for India (MIDAS) initiative as a viable model to transition LMICs from "data poverty" to "data sovereignty." MIDAS uses a rigorous, 4-domain Dataset Quality Matrix to ensure representativeness, documentation, technical fidelity, and governance, thereby creating openly benchmarked, gold-standard datasets tailored to local contexts. Initial releases, including datasets for oral and dural lesions, demonstrate the feasibility and practical value of developing robust, generalizable AI models., Furthermore, we propose a multilateral South-South Data Commons structured around 3 foundational pillars: a harmonized dataset-grading rubric, distributed custodial governance, and outcome-linked incentives. This infrastructure supports local stewardship, encourages global collaboration, and ensures that quality benchmarks drive financial incentives for dataset expansion and diversity., This proposed framework not only positions LMICs as autonomous data stewards but also enhances global AI equity. By institutionalizing quality control, interoperability, and outcome accountability, LMICs can transform from passive data consumers into active contributors to essential, trustworthy, and globally relevant biomedical datasets.