Chang Liu, Yuwei Luo, Tianzi Chen, Jie You, Jinping Wang
In this benchmark, F1, emitted-detection calibration, count behavior, and maximum observed-domain D-ECE1 were not interchangeable endpoints. The tested domain-adversarial configuration did not establish a directionally stable calibration benefit. Evaluations should report detection-specific calibration, retained-detection denominators, count consequences, and finite-benchmark observed-domain D-ECE1 alongside F1.
BACKGROUND: Traditional PCOS assessment relies on the integration of clinical history, biochemical evaluation, and pelvic ultrasonography, which may limit the efficiency of preliminary screening in some clinical settings. This study aimed to develop an AI-assisted multimodal screening model by integrating tongue-image phenotypes with structured clinical, endocrine, and metabolic variables for PCOS risk assessment.
METHODS: We collected a multimodal dataset comprising standardized tongue images and corresponding demographic, anthropometric, reproductive, endocrine, and metabolic variables. First, a pre-trained deep learning model was utilized to extract objective tongue features, including color, texture, and coating characteristics. Subsequently, a cross-modal feature fusion framework was designed to integrate these digital tongue phenotypes with clinical data. Machine learning classifiers were employed to construct the binary classification model.
RESULTS: The experimental results demonstrated that the proposed cross-modal fusion model outperformed single-modal approaches (using either tongue images or clinical data alone). The model achieved an Area Under the Curve (AUC) of 89.20%, and a Sensitivity of 85.96%. Feature importance analysis revealed that specific tongue color parameters and menstrual cycle irregularities were the most significant predictors, consistent with clinical findings.
CONCLUSION: The integration of tongue-image phenotypes with structured clinical, endocrine, and metabolic variables showed promising performance for PCOS screening in the internally validated cohort. The proposed model should be considered an adjunctive screening and risk-stratification tool rather than a replacement for routine clinical assessment and established diagnostic procedures.