EDBT 2026 Demo / reviewers in the wild / expert
Jiayi Zhang 0006
dblp:23/8208-6
· DBLP profile ↗
6ranked-venue papers
2as first author
6since 2021 · last 2026
0009-0007-5023-3952ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 3 since 2021Computer networks · 3 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Boosting Gradient-Based Training Diagnosis for Efficient and Accurate Federated LearningabstractFederated Learning (FL) allows edge clients to collaborate in model training with data privacy preserved, yet it is known to suffer low training efficiency and model accuracy. Given that efficiency and accuracy are usually conflicting objectives, existing practices increasingly employ an adaptive scheme that changes the FL configurations (e.g., quantization or sparsification level) based on runtime training status, for which accurate training diagnosis—used for guiding the optimization actions—is crucial. However, while training diagnosis is a common task shared by different optimization schemes, existing works propose their diagnosis methods in an ad-hoc manner, which yield multiple limitations. First, the diagnosis metric in an optimization scheme may sometimes be less accurate than others; second, existing schemes fail to fully exploit the diagnosis result by applying it for only one optimization action; third, existing methods usually do not perceive cross-client data heterogeneity, failing to simultaneously enhance FL accuracy. To tackle those limitations, we make a systematical study on the training diagnosis methods of multiple optimization schemes, and propose metric grafting—replacing a scheme's diagnosis metric with a better one to improve the training performance. Moreover, to fully exploit the potential of training diagnosis, we build a system platform that supports flexible combinations of training diagnosis and optimization actions (i.e., single-diagnosis-multiple actions and multiple-diagnosis-multiple-actions). Evaluation on testbeds show that, with metric grafting and advanced diagnosis action combinations, we can substantially improve the efficiency and accuracy performance of FL. Jiayi Zhang 0006, Zuo Gan, Chen Chen 0067, Zhifeng Jiang 0001, Hao Wang 0022, Yifei Zhu 0001, Quan Chen 0002, Minyi Guo |
IEEE Trans. Mob. Comput. | 1 |
| 2026 | Mitigating Server-Side Communication Bottlenecks in Distributed Learning With Round-Robin Participant CoordinationabstractDeep neural networks are increasingly trained in a distributed manner—in either clusters or with federated devices, where the participants jointly refine the global model with their gradients calculated locally. More often than not, those gradients are collected to a central server in a synchronous manner to avoid the negative impact of stale updates. However, when all the participants communicate their gradients to the server in such a uniform pace, the network on the server side—under intense contention—often becomes a performance bottleneck. To address this problem, for the cluster environment we propose theRound-Robin Synchronous Parallel(R2SP) scheme, which coordinates the participants to make updates in anevenly-gapped,round-robinmanner. This way, we can minimize the network contention with a minimum cost of the update quality; we also propose to incorporate adaptive batch sizing in R2SP to address the hardware heterogeneity among workers. Moreover, for the federated learning (FL) scenarios, we note that it is necessary yet challenging to apply the insight of R2SP to mitigate the network bottleneck in the FL server, given that there are a huge number of participants with unstable resources and inconsistent data distributions. To tackle those challenges, we further propose FL-R2SP, which extents the coordination units from individual participants to participantgroups—with the resource instability and data heterogeneity tackled within each group. We have implemented R2SP and FL-R2SP respectively with TensorFlow and PyTorch, and extensive EC2 experiments show that R2SP and FL-R2SP can respectively speed up model convergence for clustered and federated scenarios by over 20%. Jiayi Zhang 0006, Chen Chen 0067, Zuo Gan, Wei Wang 0030, Bo Li 0001, Minyi Guo |
IEEE Trans. Netw. | 1 |
| 2026 | Flexible Synchronization Control for Accurate and Efficient Federated LearningabstractFederated Learning (FL) is a distributed paradigm that supports collaborated model training while preserving data privacy, where clients periodically synchronize their local gradients once after multiple local iterations. Due to non-uniform data distribution and poor network condition, FL processes often suffer degraded training accuracy and efficiency. In this work, we analyze the microscopic parameter variation behaviors in FL, and find that an effective method to improve FL accuracy is to switch to more frequent synchronization at proper moments. In particular, such frequency-tuning moments—which can be detected from gradient characteristics—areheterogeneousacross different parameters. Motivated by such observations, we propose Parameter-Adaptive Synchronization (PAS), a FL scheme that adaptively tunes the synchronization period for each scalar parameter. The benefits of PAS are two-fold: By switching to more frequent synchronization when necessary, we can improve the FL training accuracy; by synchronizing different parameters independently, we can enable communication-computation overlapping and enhance the network utilization. We have theoretically demonstrated the convergence validity of PAS, and have further extended it with adaptive sparsification capability to jointly reduce the overall communication volume. We implemented PAS atop PyTorch, and extensive experiments show that it can substantially improve FL performance in both accuracy and communication efficiency. Zuo Gan, Chen Chen 0067, Jiayi Zhang 0006, Yifei Zhu 0001, Jieru Zhao, Quan Chen 0002, Minyi Guo |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2026 | Enabling Client-Autonomous Training Optimizations for Efficient Federated LearningabstractFederated Learning (FL) enables collaborate model training without privacy violation, where clients periodically report their updates to the server in communication rounds. Due to heterogeneous resource and limited bandwidth, FL processes often suffer from low efficiency. Existing works in that regard are oblivious to the intra-round execution status on clients, failing to tackle runtime stragglers or hide the communication overheads for some early-converged layers. In this paper, we propose FedCA, a novel mechanism that allows clients to autonomously exploit intra-round training status for higher efficiency while preserving accuracy performance. We first devise a metric to help quantify the statistical contribution of different iterations in a round, which can be efficiently profiled at runtime with the periodical sampling strategy. With the instantaneous system and statistical status, to improve computation efficiency, clients under FedCA can adaptively determine the intra-round workloads based on a utility function depicting the marginal computation benefit. Besides, to mitigate the communication bottleneck, for some parameters attaining fast local convergence, clients under FedCA can eagerly transmit their updates to the FL server prior to round completion. We also extend FedCA to FedCA+, integrating speculative sparsification to futher reduce the cumulative communication amount. We implemented FedCA and FedCA+ atop PyTorch, and large-scale experiments show that they can improve the FL efficiency by up to 45.3%. Na Lyu, Jiayi Zhang 0006, Zhi Shen, Chen Chen 0067, Zhifeng Jiang 0001, Quan Chen 0002, Minyi Guo |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2024 | FedCA: Efficient Federated Learning with Client AutonomyabstractFederated Learning (FL) enables collaborate model training without privacy violation, where clients periodically report their updates to the server in communication rounds. Due to heterogeneous resource and limited bandwidth, FL processes often suffer from low efficiency. Existing works in that regard are oblivious to the intra-round execution status on clients; however, such status information has great potential to support flexible efficiency optimizations. In this paper, we propose FedCA, a novel mechanism that allows clients to autonomously exploit intra-round training status for higher efficiency. We first devise a metric to help quantify the statistical contribution of different iterations in a round, which can be efficiently profiled at runtime with the periodical sampling strategy. With the instantaneous system and statistical status, to improve computation efficiency, clients under FedCA can adaptively determine the intra-round workloads based on a utility function. Besides, to mitigate the communication bottleneck, for some parameters attaining fast local convergence, clients under FedCA can eagerly transmit their updates to the FL server prior to round completion. We implemented FedCA atop PyTorch, and large-scale experiments show that it can improve FL efficiency by over 15%. Zhi Shen, Chen Chen 0067, Zhifeng Jiang 0001, Jiayi Zhang 0006, Quan Chen 0002, Minyi Guo |
ICPP | 5 |
| 2024 | PAS: Towards Accurate and Efficient Federated Learning with Parameter-Adaptive SynchronizationabstractFederated Learning (FL) is a distributed paradigm that supports collaborated model training while preserving data privacy, where clients periodically synchronize their local gradients once after multiple local iterations. Due to non-uniform data distribution and poor network condition, FL processes often suffer degraded training accuracy and efficiency. In this work, we analyze the microscopic parameter variation behaviors in FL, and find that an effective method to improve FL accuracy is to switch to more frequent synchronization at proper moments. Moreover, such moments can be detected from gradient characteristics, and are heterogeneous across different parameters. Motivated by such observations, we propose Parameter-Adaptive Synchronization (PAS), a FL scheme that adaptively tunes the synchronization period for each scalar parameter. The benefits of PAS are two-fold: By switching to more frequent synchronization when necessary, we can improve the FL training accuracy; by synchronizing different parameters independently, we can enable communication-computation overlapping and enhance the network utilization. We implemented PAS atop PyTorch, and extensive experiments show that it can substantially improve FL performance in both accuracy and communication efficiency. Zuo Gan, Chen Chen 0067, Jiayi Zhang 0006, Gaoxiong Zeng, Yifei Zhu 0001, Jieru Zhao, Quan Chen 0002, Minyi Guo |
IWQoS | 3 |