Filip Winzell, Ida Arvidsson, Niels Christian Overgaard, Anders Heyden, Kalle Åström, Linda Karlsson, Jacob W Vogel, Oskar Hansson, Niklas Mattsson-Carlgren, Alzheimer's Disease Neuroimaging Initiative
Methods with differential privacy achieved high privacy ratings but low levels of utility. Deep learning methods like Tabular Prior-data Fitted Network (TabPFN) and Conditional Generative Adversarial Network (CTGAN) also showed high privacy with limited utility. In contrast, non-private DataSynthesizer and Synthpop offered higher utility at a cost of lower privacy.
INTRODUCTION: The scarcity of large, clinically relevant cohorts is becoming a bottleneck in Alzheimer's disease (AD) research, as their sensitive nature makes open data sharing difficult. Privacy-preserving synthetic datasets generated with machine learning may help address this challenge.
METHODS: We compared five frameworks for generating synthetic tabular data from the Alzheimer's Disease Neuroimaging Initiative and Anti-Amyloid Treatment in Asymptomatic Alzheimer's Disease cohorts, with a set of empirical privacy and utility metrics. Two of the methods, DataSynthesizer and TableDiffusion, provide ε -differential privacy guarantees.
RESULTS: Methods with differential privacy achieved high privacy ratings but low levels of utility. Deep learning methods like Tabular Prior-data Fitted Network (TabPFN) and Conditional Generative Adversarial Network (CTGAN) also showed high privacy with limited utility. In contrast, non-private DataSynthesizer and Synthpop offered higher utility at a cost of lower privacy.
DISCUSSION: The evaluated methods demonstrated a clear trade-off between privacy and utility. High privacy was generally associated with insufficient utility, highlighting the need for further research into synthetic data generation for AD.