VLDB 2026 Research / reviewers in the wild / expert
Zihan Chen 0001
dblp:139/3503-1
· DBLP profile ↗
16ranked-venue papers
3as first author
16since 2021 · last 2026
0000-0003-0814-3391ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 7 · 7 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Global Adaptive Momentum Meets Local Personalized Perturbation: Efficient Federated LLM Fine-Tuning with Zeroth-Order GradientsabstractFederated fine-tuning of large language models (LLMs) provides a privacy-preserving approach to deploying pervasive generative AI services, yet the substantial memory overhead of first-order (FO) gradient computation presents significant practical challenges.While zeroth-order (ZO) optimization methods offer memory-efficient alternatives, they remain susceptible to performance degradation brought by data heterogeneity.Specifically, direct ZO-for-FO substitution is incompatible with existing strategies tailored for cross-client discrepancies.In response, we propose a new federated LLM fine-tuning framework, with a holistic revamped design of the entire ZO gradient processing pipeline.Crucially, with our proposed global adaptive optimization and local personalized perturbation, we present a unified solution for incorporating ZO gradients in federated learning, from local personalized perturbation sampling and ZO gradient transmission, to global ZO gradient reconstruction and aggregation with adaptive momentum, thereby directly addressing the challenges of inefficiencies and cross-client discrepancies.Our convergence analysis and experimental results demonstrate the superiority of our proposed framework over diverse heterogeneous data settings, both in terms of generalization and efficiency. Zihan Chen 0001, Howard H. Yang, Tony Q. S. Quek, Kai Fong Ernest Chong |
ACL (1) | 1 |
| 2026 | Heterogeneity-aware high-efficiency federated learning with hybrid synchronous-asynchronous splitting strategy
Zijian Li 0007, Kunyu Zhang, Bingcai Wei, Hongbo Liu 0001, Zihan Chen 0001, Xinqiang Xie, Tony Q. S. Quek |
Neural Networks | 6 |
| 2026 | Diffusion-Enabled Secure Semantic Communication Against EavesdroppingabstractThis paper proposes a novel diffusion-enabled pluggable encryption/decryption modules design against semantic eavesdropping, where the pluggable modules are optionally assembled into the semantic communication system for preventing eavesdropping. Inspired by the artificial noise (AN)-based security schemes in traditional wireless communication systems, in this paper, AN is introduced into semantic communication systems to prevent semantic eavesdropping. However, the introduction of AN also poses challenges for the legitimate receiver in extracting semantic information. Recently, denoising diffusion probabilistic models (DDPM) have demonstrated their powerful capabilities in generating multimedia content. Here, the paired pluggable modules are carefully designed using DDPM. Specifically, the pluggable encryption module generates AN and adds it to the output of the semantic transmitter, while the pluggable decryption module before semantic receiver uses DDPM to generate the detailed semantic information by removing both AN and the channel noise. In the scenario where the transmitter lacks eavesdropper’s knowledge, the artificial Gaussian noise (AGN) is used as AN. We first model a power allocation optimization problem to determine the power of AGN, in which the objective is to minimize the weighted sum of data reconstruction error of legal link, the mutual information of illegal link, and the channel input distortion. Then, a deep reinforcement learning framework using deep deterministic policy gradient is proposed to solve the optimization problem. In the scenario where the transmitter is aware of the eavesdropper’s knowledge, we propose an AN generation method based on adversarial residual networks (ARN). Unlike the previous scenario, the mutual information term in the objective function is replaced by the confidence of eavesdropper correctly retrieving private information. The adversarial residual network is then trained to minimize the modified objective function. Simulation results show that the diffusion-enabled pluggable encryption module prevents semantic eavesdropping with high covertness while the pluggable decryption module achieves the high-quality semantic communication. Boxiang He, Zihan Chen 0001, Fanggang Wang 0001, Shilian Wang, Zhijin Qin, Tony Q. S. Quek |
IEEE Trans. Wirel. Commun. | 2 |
| 2025 | Robust Federated Learning Over the Air: Combating Heavy-Tailed Noise with Median Anchored ClippingabstractLeveraging over-the-air computations for model aggregation is an effective approach to cope with the communication bottleneck in federated edge learning. By exploiting the superposition properties of multi-access channels, this approach facilitates an integrated design of communication and computation, thereby enhancing system privacy while reducing implementation costs. However, the inherent electromagnetic interference in radio channels often exhibits heavy-tailed distributions, giving rise to exceptionally strong noise in globally aggregated gradients that can significantly deteriorate the training performance. To address this issue, we propose a novel gradient clipping method, termed Median Anchored Clipping (MAC), to combat the detrimental effects of heavy-tailed noise. We also derive analytical expressions for the convergence rate of model training with analog over-the-air federated learning under MAC, which quantitatively demonstrates the effect of MAC on training performance. Extensive experimental results show that the proposed MAC algorithm effectively mitigates the impact of heavy-tailed noise, hence substantially enhancing system robustness. Zihan Chen 0001, Kai Fong Ernest Chong, Bikramjit Das, Tony Q. S. Quek, Howard H. Yang |
WiOpt | 2 |
| 2025 | Model-Heterogeneous Prototypical Federated Learning Over the AirabstractOver-the-air federated learning (OTA FL) provides a joint computation and communication approach to design FL systems with improved efficiency. By leveraging the superposition property of wireless channels, OTA FL enables the automatic aggregation of intermediate parameters-such as gradients-across a large number of clients, significantly reducing communication overhead while concurrently enhancing transmission privacy. However, gradient aggregation requires all clients to use identical model architectures, a condition often impractical in real-world scenarios due to variations in client hardware and computational capabilities. This mismatch can lead to scalability issues and system incompatibilities. To address this challenge, we propose a model-agnostic method based on model prototypes that enables collaborative training across clients with heterogeneous models. The proposed method bypasses the requirements of the conventional model weight/gradient updates with prototype vector aggregation, without requiring the model structures of all clients to be identical. To the best of our knowledge, the proposed method is the first to explore the prototypical model-heterogeneous OTA FL with desirable training performance and extremely low communication cost. We conducted extensive experiments to verify the efficacy of the proposed method. The results show that our approach not only significantly reduces communication overhead but also exploits the superior capabilities of large models to enhance the performance of smaller models. Chuhan Sun, Zihan Chen 0001, Liyinglan Liu, Tony Q. S. Quek, Howard H. Yang |
WiOpt | 2 |
| 2025 | Sparsified Random Partial Model Update for Personalized Federated LearningabstractFederated Learning (FL) stands as a privacy-preserving machine learning paradigm that enables collaborative training of a global model across multiple clients. However, the practical implementation of FL models often confronts challenges arising from data heterogeneity and limited communication resources. To address the aforementioned issues simultaneously, we develop a Sparsified Random Partial Update framework for personalized Federated Learning (SRP-pFed), which builds upon the foundation of dynamic partial model updates. Specifically, we decouple the local model into personal and shared parts to achieve personalization. For each client, the ratio of its personal part associated with the local model, referred to as the update rate, is regularly renewed over the training procedure via a random walk process endowed with reinforced memory. In each global iteration, clients are clustered into different groups where the ones in the same group share a common update rate. Benefiting from such design,SRP-pFedrealizes model personalization while substantially reducing communication costs in the uplink transmissions. We conduct extensive experiments on various training tasks with diverse heterogeneous data settings. The results demonstrate that theSRP-pFedconsistently outperforms the state-of-the-art methods in test accuracy and communication efficiency. Zihan Chen 0001, Chenyuan Feng, Geyong Min, Tony Q. S. Quek, Howard H. Yang |
IEEE Trans. Mob. Comput. | 2 |
| 2025 | Personalized Federated Learning Over the AirabstractWe propose an effective approach toward implementing personalized federated learning at the edge of wireless networks. The scheme employs a bi-level optimization framework to personalize the federated learning models, and leverages over-the-air computations for model aggregation.We identify a mutual benefit in such a design. Specifically, personalized federated learning models address the challenge of data heterogeneity in federated learning, while over-the-air computations, which capitalize on the superposition property of multiple access channels, enable all clients to upload their intermediate parameters in each communication round for global aggregation, significantly enhancing system scalability. However, the channel fading and heavy-tailed noise introduced by over-the-air computations pose challenges to the robustness of personalized federated learning models. By adopting a bi-level optimization framework, we improve the stability of personalized federated learning models based on over-the-air computations, establishing a scalable and robust federated edge learning system. We also derive convergence rates of the proposed algorithm, encompassing key factors such as model compression, channel fading, and heavy-tailed noise. The analysis offers a comprehensive understanding of how system configurations affect training performance. We corroborate the efficacy of our framework via extensive experiments. Zeshen Li, Zihan Chen 0001, Tony Q. S. Quek, Howard H. Yang |
IEEE Trans. Wirel. Commun. | 2 |
| 2024 | FedLoGe: Joint Local and Generic Federated Learning under Long-tailed DataabstractFederated Long-Tailed Learning (Fed-LT), a paradigm wherein data collected from decentralized local clients manifests a globally prevalent long-tailed distribution, has garnered considerable attention in recent times. In the context of Fed-LT, existing works have predominantly centered on addressing the data imbalance issue to enhance the efficacy of the generic global model while neglecting the performance at the local level. In contrast, conventional Personalized Federated Learning (pFL) techniques are primarily devised to optimize personalized local models under the presumption of a balanced global data distribution. This paper introduces an approach termed Federated Local and Generic Model Training in Fed-LT (FedLoGe), which enhances both local and generic model performance through the integration of representation learning and classifier alignment within a neural collapse framework. Our investigation reveals the feasibility of employing a shared backbone as a foundational framework for capturing overarching global trends, while concurrently employing individualized classifiers to encapsulate distinct refinements stemming from each client’s local features. Building upon this discovery, we establish the Static Sparse Equiangular Tight Frame Classifier (SSE-C), inspired by neural collapse principles that naturally prune extraneous noisy features and foster the acquisition of potent data representations. Furthermore, leveraging insights from imbalance neural collapse's classifier norm patterns, we develop Global and Local Adaptive Feature Realignment (GLA-FR) via an auxiliary global classifier and personalized Euclidean norm transfer to align global features with client preferences. Extensive experimental results on CIFAR-10/100-LT, ImageNet, and iNaturalist demonstrate the advantage of our method over state-of-the-art pFL and Fed-LT approaches. Zikai Xiao, Zihan Chen 0001, Liyinglan Liu, Yang Feng 0011, Joey Tianyi Zhou, Jian Wu 0001, Wanlu Liu, Howard H. Yang, Zuozhu Liu |
ICLR | 2 |
| 2024 | Exploiting Complex Network-Based Clustering for Personalization-Enhanced Hierarchical Federated Edge LearningabstractFederated Learning (FL) has been extensively applied in urban environmental prediction tasks of mobile edge computing by training a global machine learning model without data sharing. However, the training of FL faces the challenges such as the poor generalization capability of a single global model over heterogeneous data and hefty communication overhead caused by the frequent model exchange between massive edge servers and remote cloud servers. To address such issues, we propose HPFL-CN, a novel communication-efficient Hierarchical Personalized Federated edge Learning framework with Complex Network clustering. HPFL-CN introduces Privacy-preserving Feature Clustering (PFC) to extract privacy-preserving low-dimensional feature representations of each edge server via mapping the environmental data to different complex network domains for clustering similar edge servers accurately. Based on the clustering results of PFC, anedge-mediator-cloudhierarchical architecture is proposed to realize personalization at the cluster level by Effective Hierarchical Scheduling (EHS). Furthermore, to adapt to dynamic scenarios of new edge servers joining and streaming data generation, we further extend HPFL-CN to Adaptive personalized federated learning with dynamic grouping (Ada-HPFL-CN), which can flexibly re-group edge servers and adjust mixed model weights and the model aggregation frequency adaptively. Our extensive experiments on real-world datasets demonstrate the efficacy of our framework, which outperforms state-of-the-art FL methods regarding personalization and communication efficiency performance. Zijian Li 0007, Zihan Chen 0001, Xiaohui Wei 0002, Shang Gao 0005, Hengshan Yue, Zhewen Xu, Tony Q. S. Quek |
IEEE Trans. Mob. Comput. | 2 |
| 2023 | Personalizing Federated Learning with Over-The-Air ComputationsabstractFederated edge learning is a promising technology to deploy intelligence at the edge of wireless networks in a privacy-preserving manner. Under such a setting, multiple clients collaboratively train a global generic model under the coordination of an edge server. But the training efficiency is often hindered by challenges arising from limited communication and data heterogeneity. In this paper, we present a distributed training paradigm that employs analog over-the-air computation to alleviate the communication bottleneck. Additionally, we leverage a bi-level optimization framework to personalize the federated learning model so as to cope with the data heterogeneity issue. As a result, it enhances the generalization and robustness of each client’s local model. We elaborate on the model training procedure and its advantages over conventional frameworks. We provide a convergence analysis that theoretically demonstrates the training efficiency. We also conduct extensive experiments to validate the efficacy of the proposed framework. Zihan Chen 0001, Zeshen Li, Howard H. Yang, Tony Q. S. Quek |
ICASSP | 1 |
| 2023 | Spectral Co-Distillation for Personalized Federated LearningabstractPersonalized federated learning (PFL) has been widely investigated to address the challenge of data heterogeneity, especially when a single generic model is inadequate in satisfying the diverse performance requirements of local clients simultaneously. Existing PFL methods are inherently based on the idea that the relations between the generic global and personalized local models are captured by the similarity of model weights. Such a similarity is primarily based on either partitioning the model architecture into generic versus personalized components or modeling client relationships via model weights. To better capture similar (yet distinct) generic versus personalized model representations, we propose $\textit{spectral distillation}$, a novel distillation method based on model spectrum information. Building upon spectral distillation, we also introduce a co-distillation framework that establishes a two-way bridge between generic and personalized model training. Moreover, to utilize the local idle time in conventional PFL, we propose a wait-free local training protocol. Through extensive experiments on multiple datasets over diverse heterogeneous data settings, we demonstrate the outperformance and efficacy of our proposed spectral co-distillation method, as well as our wait-free training protocol. Zihan Chen 0001, Howard H. Yang, Tony Q. S. Quek, Kai Fong Ernest Chong |
NeurIPS | 1 |
| 2023 | Fed-GraB: Federated Long-tailed Learning with Self-Adjusting Gradient BalancerabstractData privacy and long-tailed distribution are the norms rather than the exception in many real-world tasks. This paper investigates a federated long-tailed learning (Fed-LT) task in which each client holds a locally heterogeneous dataset; if the datasets can be globally aggregated, they jointly exhibit a long-tailed distribution. Under such a setting, existing federated optimization and/or centralized long-tailed learning methods hardly apply due to challenges in (a) characterizing the global long-tailed distribution under privacy constraints and (b) adjusting the local learning strategy to cope with the head-tail imbalance. In response, we propose a method termed $\texttt{Fed-GraB}$, comprised of a Self-adjusting Gradient Balancer (SGB) module that re-weights clients' gradients in a closed-loop manner, based on the feedback of global long-tailed distribution evaluated by a Direct Prior Analyzer (DPA) module. Using $\texttt{Fed-GraB}$, clients can effectively alleviate the distribution drift caused by data heterogeneity during the model training process and obtain a global model with better performance on the minority classes while maintaining the performance of the majority classes. Extensive experiments demonstrate that $\texttt{Fed-GraB}$ achieves state-of-the-art performance on representative datasets such as CIFAR-10-LT, CIFAR-100-LT, ImageNet-LT, and iNaturalist. Zikai Xiao, Zihan Chen 0001, Songshang Liu, Hualiang Wang, Yang Feng 0011, Jin Hao, Joey Tianyi Zhou, Jian Wu 0001, Howard H. Yang, Zuozhu Liu |
NeurIPS | 2 |
| 2023 | Federated Learning with Partial Gradients Over-the-AirabstractWe develop a theoretical framework to study the training of federated learning models with partial gradients via over-the-air computing. The system consists of an edge server and multiple clients, aiming to collaboratively minimize a global loss function. The clients conduct local training and upload the intermediate parameters (e.g. the gradients) by analog transmissions. Specifically, each client modulates the entries of its local gradient onto a set of common orthogonal waveforms and sends out the signal simultaneously to the edge server; owing to the limited number of orthogonal waveforms, only a subset of the parameters can be selected for uploading during each round of communication. On the server side, it passes the received analog signal to a bank of match filters and obtains a noisy partial gradient vector. The server then uses this partial gradient to update the global parameter and feeds the new model back to all the clients for another round of local training. We derive the convergence rate of such a model training algorithm. We also conduct experiments to investigate the effects of different masking schemes on the convergence performance. The findings advance the understanding of over-the-air federated learning and provide useful insights for system designs. Wendi Wang 0005, Zihan Chen 0001, Nikolaos Pappas 0001, Howard H. Yang |
SECON | 2 |
| 2022 | FedCorr: Multi-Stage Federated Learning for Label Noise CorrectionabstractFederated learning (FL) is a privacy-preserving distributed learning paradigm that enables clients to jointly train a global model. In real-world FL implementations, client data could have label noise, and different clients could have vastly different label noise levels. Although there exist methods in centralized learning for tackling label noise, such methods do not perform well on heterogeneous label noise in FL settings, due to the typically smaller sizes of client datasets and data privacy requirements in FL. In this paper, we propose FedCorr, a general multi-stage framework to tackle heterogeneous label noise in FL, without making any assumptions on the noise models of local clients, while still maintaining client data privacy. In particular, (1) FedCorr dynamically identifies noisy clients by exploiting the dimensionalities of the model prediction subspaces independently measured on all clients, and then identifies incorrect labels on noisy clients based on persample losses. To deal with data heterogeneity and to increase training stability, we propose an adaptive local proximal regularization term that is based on estimated local noise levels. (2) We further finetune the global model on identified clean clients and correct the noisy labels for the remaining noisy clients after finetuning. (3) Finally, we apply the usual training on all clients to make full use of all local data. Experiments conducted on CIFAR-10/100 with federated synthetic label noise, and on a real-world noisy dataset, Clothing1M, demonstrate that FedCorr is robust to label noise and substantially outperforms the state-of-the-art methods at multiple noise levels. Zihan Chen 0001, Tony Q. S. Quek, Kai Fong Ernest Chong |
CVPR | 2 |
| 2022 | HPFL-CN: Communication-Efficient Hierarchical Personalized Federated Edge Learning via Complex Network Feature ClusteringabstractFederated Learning (FL), a promising privacy-preserving distributed learning paradigm, has been extensively applied in urban environmental prediction tasks of Mobile Edge Computing (MEC) by training a global machine learning model without data sharing. However, it is hard for the shared global model to be well generalized among local edge servers, due to the statistical data heterogeneity, especially in real-world urban environmental data. Besides, the existing FL approaches may result in excessive communication and computation overhead due to the frequent transmission and aggregation of model parameters between massive edge servers and remote cloud servers. To address the above issues, we propose HPFL-CN, a novel communication-efficient Hierarchical Personalized Federated edge Learning framework via Complex Network feature clustering, aiming to cluster edge servers with similar environmental data distributions and then high-efficiently train personalized models for each cluster via hierarchical architecture. Specifically, HPFL-CN introduces Privacy-preserving Feature Clustering (PFC) to extract privacy-preserving low-dimensional feature representations of each edge server via mapping the environmental data to different complex network domains for clustering similar edge servers accurately. According to the clustering results of PFC, HPFL-CN further introduces an edge-mediator-cloud architecture for hierarchical model aggregation by Effective Hierarchical Scheduling (EHS), in which every mediator coordinates the training of edge servers within each cluster and periodically uploads model to cloud server for global model aggregation. Meanwhile, each mediator server would find a trade-off between cloud and edge models to realize personalization within clusters. Our extensive experiments on real-world datasets demonstrate the effectiveness and generalization of HPFL-CN, which outperforms other state-of-the-art FL methods regarding personalization performance and communication efficiency. Zijian Li 0007, Zihan Chen 0001, Xiaohui Wei 0002, Shang Gao 0005, Chenghao Ren, Tony Q. S. Quek |
SECON | 2 |
| 2022 | Joint Scheduling and Resource Allocation for Hierarchical Federated Edge LearningabstractThe concept of hierarchical federated edge learning (H-FEEL) has been recently proposed as an enhancement of federated learning model. Such a system generally consists of three entities, i.e., the server, helpers, and clients, in which each helper collects the trained gradients from clients nearby, aggregates them, and sends the result to the server for global model update. Due to limited communication resources, only a portion of helpers can be scheduled to upload their aggregated gradients in each round of the model training. And that necessitates a well-designed scheme for the joint helper scheduling and communication resources allocation. In this paper, we develop a training algorithm for the H-FEEL system which involves local gradient computing, weighted gradient uploading, and machine learning model updating phases. By characterizing these phases mathematically and analyzing one-round convergence bound of the training algorithm, we formulate an optimization problem to achieve the scheduling and resource allocation scheme. The problem simultaneously captures the uncertainty of the wireless channel and the importance of the weighted gradient. To solve the problem, we first transform it into an equivalent problem and then decompose the transformed problem into two subproblems:bit and sub-channel allocationandhelper scheduling, which are mixed integer nonlinear programming and continuous nonlinear problems, respectively. For the first subproblem, we obtain an optimal solution of exponential complexity and a suboptimal solution that has polynomial complexity. For the second subproblem, we obtain a closed-form optimal solution in a special case and a suboptimal solution in the general case. The efficacy of our scheme is amply demonstrated via simulations and the analytical framework is shown to provide valuable design insights for the practical implementation of the H-FEEL system. Wanli Wen, Zihan Chen 0001, Howard H. Yang, Wenchao Xia, Tony Q. S. Quek |
IEEE Trans. Wirel. Commun. | 2 |