EDBT 2026 Demo / reviewers in the wild / expert
Zhifeng Jiang 0001
dblp:82/4796-1
· DBLP profile ↗
10ranked-venue papers
4as first author
10since 2021 · last 2026
0000-0002-7024-441XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 2 first-author · 6 since 2021Security and privacy · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Boosting Gradient-Based Training Diagnosis for Efficient and Accurate Federated LearningabstractFederated Learning (FL) allows edge clients to collaborate in model training with data privacy preserved, yet it is known to suffer low training efficiency and model accuracy. Given that efficiency and accuracy are usually conflicting objectives, existing practices increasingly employ an adaptive scheme that changes the FL configurations (e.g., quantization or sparsification level) based on runtime training status, for which accurate training diagnosis—used for guiding the optimization actions—is crucial. However, while training diagnosis is a common task shared by different optimization schemes, existing works propose their diagnosis methods in an ad-hoc manner, which yield multiple limitations. First, the diagnosis metric in an optimization scheme may sometimes be less accurate than others; second, existing schemes fail to fully exploit the diagnosis result by applying it for only one optimization action; third, existing methods usually do not perceive cross-client data heterogeneity, failing to simultaneously enhance FL accuracy. To tackle those limitations, we make a systematical study on the training diagnosis methods of multiple optimization schemes, and propose metric grafting—replacing a scheme's diagnosis metric with a better one to improve the training performance. Moreover, to fully exploit the potential of training diagnosis, we build a system platform that supports flexible combinations of training diagnosis and optimization actions (i.e., single-diagnosis-multiple actions and multiple-diagnosis-multiple-actions). Evaluation on testbeds show that, with metric grafting and advanced diagnosis action combinations, we can substantially improve the efficiency and accuracy performance of FL. Jiayi Zhang 0006, Zuo Gan, Chen Chen 0067, Zhifeng Jiang 0001, Hao Wang 0022, Yifei Zhu 0001, Quan Chen 0002, Minyi Guo |
IEEE Trans. Mob. Comput. | 4 |
| 2026 | Enabling Client-Autonomous Training Optimizations for Efficient Federated LearningabstractFederated Learning (FL) enables collaborate model training without privacy violation, where clients periodically report their updates to the server in communication rounds. Due to heterogeneous resource and limited bandwidth, FL processes often suffer from low efficiency. Existing works in that regard are oblivious to the intra-round execution status on clients, failing to tackle runtime stragglers or hide the communication overheads for some early-converged layers. In this paper, we propose FedCA, a novel mechanism that allows clients to autonomously exploit intra-round training status for higher efficiency while preserving accuracy performance. We first devise a metric to help quantify the statistical contribution of different iterations in a round, which can be efficiently profiled at runtime with the periodical sampling strategy. With the instantaneous system and statistical status, to improve computation efficiency, clients under FedCA can adaptively determine the intra-round workloads based on a utility function depicting the marginal computation benefit. Besides, to mitigate the communication bottleneck, for some parameters attaining fast local convergence, clients under FedCA can eagerly transmit their updates to the FL server prior to round completion. We also extend FedCA to FedCA+, integrating speculative sparsification to futher reduce the cumulative communication amount. We implemented FedCA and FedCA+ atop PyTorch, and large-scale experiments show that they can improve the FL efficiency by up to 45.3%. Na Lyu, Jiayi Zhang 0006, Zhi Shen, Chen Chen 0067, Zhifeng Jiang 0001, Quan Chen 0002, Minyi Guo |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2025 | SGDRC: Software-Defined Dynamic Resource Control for Concurrent DNN Inference on NVIDIA GPUsabstractCloud service providers heavily colocate high-priority, latency sensitive (LS), and low-priority, best-effort (BE) DNN inference services on the same GPU to improve resource utilization in data centers. Among the critical shared GPU resources, there has been very limited analysis on the dynamic allocation of compute units and VRAM bandwidth, mainly for two reasons: (1) The native GPU resource management solutions are either hardware-specific, or unable to dynamically allocate resources to different tenants, or both; (2) NVIDIA doesn't expose interfaces for VRAM bandwidth allocation, and the software stack and VRAM channel architectures are black-box, both of which limit the software-level resource management. These drive prior work to design either conservative sharing policies detrimental to throughput, or static resource partitioning only applicable to a few GPU models. Yongkang Zhang 0003, Haoxuan Yu, Chenxia Han, Baotong Lu, Zhifeng Jiang 0001, Yang Li 0090, Xiaowen Chu 0001, Huaicheng Li |
PPoPP | 7 |
| 2025 | Feature Reconstruction Attacks and Countermeasures of DNN Training in Vertical Federated LearningabstractFederated learning (FL) has increasingly been deployed, in its vertical form, among organizations to facilitate secure collaborative training. In vertical FL (VFL), participants hold disjoint features of the same set of sample instances. The one withlabels- theactive party, initiates training and interacts with other participants - thepassive parties. It remains largely unknownwhetherandhowan active party can extract private feature data owned by passive parties, especially when training deep neural network (DNN) models. This work examines the feature security problem of DNN training in VFL. We consider a DNN model partitioned between active and passive parties, where the passive party holds a subset of the input layer with some features of binary values. Though proved to be NP-hard. we demonstrate that, unless the feature dimension is exceedingly large, it remains feasible, both theoretically and practically, to launch a reconstruction attack with an efficient search-based algorithm that prevails over current feature protection. We propose a novel feature protection scheme by perturbing intermediate results and fabricated input features, which effectively misleads reconstruction attacks towards pre-specified random values. The evaluation shows it sustains feature reconstruction attack in various VFL applications with negligible impact on model performance. Peng Ye 0005, Zhifeng Jiang 0001, Wei Wang 0030, Bo Li 0001, Baochun Li |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2024 | Dordis: Efficient Federated Learning with Dropout-Resilient Differential PrivacyabstractFederated learning (FL) is increasingly deployed among multiple clients to train a shared model over decentralized data. To address privacy concerns, FL systems need to safeguard the clients' data from disclosure during training and control data leakage through trained models when exposed to untrusted domains. Distributed differential privacy (DP) offers an appealing solution in this regard as it achieves a balanced tradeoff between privacy and utility without a trusted server. However, existing distributed DP mechanisms are impractical in the presence of client dropout, resulting in poor privacy guarantees or degraded training accuracy. In addition, these mechanisms suffer from severe efficiency issues. Zhifeng Jiang 0001, Wei Wang 0030, Ruichuan Chen |
EuroSys | 1 |
| 2024 | FedCA: Efficient Federated Learning with Client AutonomyabstractFederated Learning (FL) enables collaborate model training without privacy violation, where clients periodically report their updates to the server in communication rounds. Due to heterogeneous resource and limited bandwidth, FL processes often suffer from low efficiency. Existing works in that regard are oblivious to the intra-round execution status on clients; however, such status information has great potential to support flexible efficiency optimizations. In this paper, we propose FedCA, a novel mechanism that allows clients to autonomously exploit intra-round training status for higher efficiency. We first devise a metric to help quantify the statistical contribution of different iterations in a round, which can be efficiently profiled at runtime with the periodical sampling strategy. With the instantaneous system and statistical status, to improve computation efficiency, clients under FedCA can adaptively determine the intra-round workloads based on a utility function. Besides, to mitigate the communication bottleneck, for some parameters attaining fast local convergence, clients under FedCA can eagerly transmit their updates to the FL server prior to round completion. We implemented FedCA atop PyTorch, and large-scale experiments show that it can improve FL efficiency by over 15%. Zhi Shen, Chen Chen 0067, Zhifeng Jiang 0001, Jiayi Zhang 0006, Quan Chen 0002, Minyi Guo |
ICPP | 4 |
| 2024 | Lotto: Secure Participant Selection against Adversarial Servers in Federated Learning
Zhifeng Jiang 0001, Peng Ye 0005, Shiqi He, Wei Wang 0030, Ruichuan Chen, Bo Li 0001 |
USENIX Security Symposium | 1 |
| 2023 | Towards Efficient Synchronous Federated Training: A Survey on System Optimization StrategiesabstractThe increasing demand for privacy-preserving collaborative learning has given rise to a new computing paradigm called federated learning (FL), in which clients collaboratively train a machine learning (ML) model without revealing their private training data. Given an acceptable level of privacy guarantee, the goal of FL is to minimize thetime-to-accuracyof model training. Compared with distributed ML in data centers, there are four distinct challenges to achieving short time-to-accuracy in FL training, namely the lack of information for optimization, the tradeoff between statistical and system utility, client heterogeneity, and large configuration space. In this paper, we survey recent works in addressing these challenges and present them following a typical training workflow through three phases: client selection, configuration, and reporting. We also review system works including measurement studies and benchmarking tools that aim to support FL developers. Zhifeng Jiang 0001, Wei Wang 0030, Bo Li 0001, Qiang Yang 0001 |
IEEE Trans. Big Data | 1 |
| 2022 | Pisces: efficient federated learning via guided asynchronous trainingabstractFederated learning (FL) is typically performed in a synchronous parallel manner, and the involvement of a slow client delays the training progress. Current FL systems employ a participant selection strategy to select fast clients with quality data in each iteration. However, this is not always possible in practice, and the selection strategy has to navigate a knotty tradeoff between the speed and the data quality. Zhifeng Jiang 0001, Wei Wang 0030, Baochun Li, Bo Li 0001 |
SoCC | 1 |
| 2021 | Gillis: Serving Large Neural Networks in Serverless Functions with Automatic Model PartitioningabstractThe increased use of deep neural networks has stimulated the growing demand for cloud-based model serving platforms. Serverless computing offers a simplified solution: users deploy models as serverless functions and let the platform handle provisioning and scaling. However, serverless functions have constrained resources in CPU and memory, making them inefficient or infeasible to serve large neural networks-which have become increasingly popular. In this paper, we present Gillis, a serverless-based model serving system that automatically partitions a large model across multiple serverless functions for faster inference and reduced memory footprint per function. Gillis employs two novel model partitioning algorithms that respectively achieve latency-optimal serving and cost-optimal serving with SLO compliance. We have implemented Gillis on three serverless platforms-AWS Lambda, Google Cloud Functions, and KNIX-with MXNet as the serving backend. Experimental evaluations against popular models show that Gillis supports serving very large neural networks, reduces the inference latency substantially, and meets various SLOs with a low serving cost. Minchen Yu, Zhifeng Jiang 0001, Hok Chun Ng, Wei Wang 0030, Ruichuan Chen, Bo Li 0001 |
ICDCS | 2 |