Bayu Wijaya Putra, M. Rudi Sanjaya, Hasnan Afif
Clinical data quality in primary healthcare settings often faces challenges such as inconsistent terminology, missing values, duplicate entries, and relational inconsistencies across clinical variables. These issues hinder the accuracy and reliability of epidemiological analyses and predictive modeling based on real-world clinical data. This study develops and evaluates a comprehensive data-cleaning pipeline using 2,354 visit-level clinical records obtained from a primary healthcare facility, representing 496 unique patients. The proposed pipeline integrates diagnosis normalization through ICD-10 mapping supported by fuzzy matching and rule-based refinement, missing-value handling, patient identity deduplication, diagnosis–therapy relational consistency evaluation, and anomaly detection using statistical approaches (IQR, Z-score, and quantile analysis) as well as clustering techniques. The results demonstrate that 319 unique raw diagnosis strings were consolidated into 19 standardized ICD-10 codes, missing values across key clinical variables were reduced by more than 70%, and relational inconsistencies between diagnoses and therapies were substantially minimized. Visualization of diagnosis–therapy relationships reveals more coherent and interpretable clinical patterns after standardization. The novelty of this study lies in the development of a multi-method, integrated cleaning framework tailored to primary healthcare data in Indonesia. The resulting dataset exhibits improved structural integrity, terminology consistency, and relational coherence, making it more reliable and suitable for downstream analytics and predictive modeling. These findings highlight that a structured data-cleaning pipeline can significantly enhance the interpretability and analytical readiness of real-world clinical data in primary care environments.