科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ JMIR medical informatics2026-09-11

Quantifying the Impact of Anonymization-Induced Clinical Data Quality Loss: Methodological Quantitative Case Study Using Primary Diagnosis Codes and Hospital Length of Stay.

Gaetan Kamdje Wabo, Piotr Pawel Sokolowski, Mahboubeh Jannesari Ladani, Michael Hagmann, Thomas Ganslandt, Fabian Siegel

一句话结论 · In one sentence

k-Anonymity preserved the ranking of diagnosis-specific LOS effects but altered distributional shape, individual effect magnitudes, and diagnostic vocabulary, none of which was flagged by the internal loss metric. Anonymized data of this type may support ordinal analyses but can mislead analyses requiring faithful variance structure, accurate absolute effects, or complete rare-diagnosis representation. A reporting checklist is provided to document these effects.

原始摘要(英文原文)· Original abstract
BACKGROUND: The secondary use of electronic health record data requires robust privacy protection. k-Anonymity is widely used to enable data sharing by ensuring that each quasi-identifier combination occurs in at least k records; yet, its analytical impact on clinically meaningful structures remains insufficiently characterized, particularly for the combination of record suppression and microaggregation that arises when a numeric attribute lacks a natural generalization hierarchy. A further gap is that anonymization tools report internal information-loss values but do not signal the downstream distributional and inferential distortions these transformations introduce. OBJECTIVE: This study evaluated the analytical footprint of k-anonymity at k=5, 10, and 15 on 2 core data elements in retrospective hospital research: primary International Classification of Diseases, 10th Revision, German Modification (ICD-10-GM) diagnosis codes, and hospital length of stay (LOS). It aimed to determine and quantify whether anonymization introduces meaningful distortions not captured by the anonymization tool itself, and whether diagnosis-specific LOS patterns remain reproducible after anonymization. METHODS: We analyzed 719,387 inpatient encounters from University Hospital Mannheim from 2010 to 2024. Anonymization was performed with the ARX tool. It used record suppression and microaggregation. Distributional distortion was assessed with the Kolmogorov-Smirnov D statistic, quantile shifts, IQR changes, and tail changes. Categorical fidelity was assessed with the Jaccard coefficient and Cramer V. Inferential reproducibility was assessed with a 3-level linear mixed model. The model included random intercepts for the three-character diagnosis codes from the International Classification of Diseases (ICD-3) and patients. We compared the intraclass correlation coefficient and diagnosis-level effect concordance. Concordance was quantified using Spearman ρ and Lin concordance correlation coefficient, both with 95% CIs. A composite traffic-light verdict summarized the results. RESULTS: ARX masked quasi-identifier cells rather than deleting rows; the proportion of encounters with a masked cell rose from 0.77% (k=5) to 2.62% (k=15), distributed almost uniformly across admission years. Kolmogorov-Smirnov D was stable at 0.147. Median LOS shifted by 1 day, and the SD declined by about 6.5 days, while the diagnosis-level mean changed considerably (absolute mean shift of -1.81 days). Jaccard overlap fell from 0.624 to 0.421. The diagnosis intraclass correlation coefficient rose from 0.294 to 0.837, reflecting variance compression rather than improved signal. Best linear unbiased prediction rank concordance (Spearman ρ 0.964-0.970) and aggregate magnitude agreement (Lin concordance correlation coefficient 0.959-0.966) were high; yet, about 7% of low-signal diagnoses showed sign reversals. Most distributional change occurred at k=5. CONCLUSIONS: k-Anonymity preserved the ranking of diagnosis-specific LOS effects but altered distributional shape, individual effect magnitudes, and diagnostic vocabulary, none of which was flagged by the internal loss metric. Anonymized data of this type may support ordinal analyses but can mislead analyses requiring faithful variance structure, accurate absolute effects, or complete rare-diagnosis representation. A reporting checklist is provided to document these effects.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Quantifying the Impact of Anonymization-Induced Clinical Data Quality Loss: Methodological Quantitative Case Study Using Primary Diagnosis Codes and Hospital Length of Stay. — 科研速览 Science Skim