科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Journal of Industrial Information Integration2025-12-27· Skew

A thorough assessment of the non-IID data impact in federated learning

Daniel M. Jimenez-Gutierrez, Mehrdad Hassanzadeh, Aris Anagnostopoulos, Ioannis Chatzigiannakis, Andrea Vitaletti

原始摘要(英文原文)· Original abstract
Federated learning (FL) allows collaborative machine learning (ML) model training among decentralized clients’ information, ensuring data privacy. The decentralized nature of FL deals with non-independent and identically distributed (non-IID) data. This open problem has notable consequences, such as decreased model performance and longer convergence times. Despite its importance, experimental studies systematically addressing all types of data heterogeneity (a.k.a. non-IIDness) remain scarce. This paper aims to fill this gap by assessing and quantifying the non-IID effect through an empirical analysis. We use the Hellinger Distance ( HD ) to measure differences in distribution among clients. Our study benchmarks five state-of-the-art strategies for handling non-IID data, including label, feature, quantity, and spatiotemporal skews, under realistic and controlled conditions. This is the first comprehensive analysis of the spatiotemporal skew effect in FL. Our findings highlight the significant impact of label and spatiotemporal skew non-IID types on FL model performance, with notable performance drops occurring at specific HD thresholds. The FL performance is also heavily affected, mainly when the non-IIDness is extreme. Thus, we provide recommendations for FL research to tackle data heterogeneity effectively. Our work represents the most extensive examination of non-IIDness in FL, offering a robust foundation for future research. • Label skew and spatiotemporal skew have the most significant impact on the model’s performance. • The drop in the model’s performance for label skew appears in a double threshold. A notable performance decline is immediately evident when the Hellinger Distance exceeds 0.5 and 0.75. • Feature skew does not alter the model’s performance nor the convergence point. • The quantity skew in the client’s data does not affect the model’s performance. • The higher the non-IIDness level in time and space (spatiotemporal skew), the worse the model’s performance.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

A thorough assessment of the non-IID data impact in federated learning — 科研速览 Science Skim