Ningxin Su, Baochun Li, Bo Li
Federated learning (FL) is a privacy-motivated paradigm for distributed training of deep learning models, as it allows a large number of clients to collaboratively train a shared global model without centralizing their private data. Since its debut with Federated Averaging as the first server aggregation algorithm, many FL algorithms have been proposed to improve performance. Yet, these algorithms are rarely benchmarked and compared under the same open-source framework and controlled configurations, and their performance claims can be difficult to substantiate in fair and reproducible studies. In this paper, we evaluate a curated and representative collection of FL algorithms in the same open-source benchmarking framework so that they can be compared fairly at scale in a reproducible fashion. To achieve this objective, we presentPlato, an open-source FL research framework that we have designed and implemented from scratch. WithPlato, we evaluate and compare algorithms spanning (1) server aggregation; (2) client training customization; (3) client selection, in both synchronous and asynchronous settings; (4) personalized federated learning; and (5) communication efficiency and payload processing. Across diverse experimental scenarios (tasks, client populations, and data distributions), we report findings and practical insights, including pitfalls and confounding factors that can lead to misleading conclusions if not reported carefully. Under our unified experimental settings and time model, Federated Averaging with random client selection remains a strong baseline and is often competitive with more complex alternatives.