EDBT 2026 Demo / reviewers in the wild / expert
Gang Yan 0002
dblp:87/7043-2
· DBLP profile ↗
15ranked-venue papers
10as first author
14since 2021 · last 2026
0000-0002-7734-1589ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 6 first-author · 7 since 2021Databases, data management, data science and information retrieval · 5 · 4 first-author · 5 since 2021Computer networks · 4 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | WPGRec: Wavelet Packet Guided Graph Enhanced Sequential Recommendation
Zhiquan Ji, Gang Yan 0002 |
SIGIR | 3 |
| 2025 | FedSTEP: Asynchronous and Staleness-Aware Personalization for Efficient Federated LearningabstractPersonalized Federated Learning (PFL) aims to provide client-specific models that adapt to local data distributions while leveraging shared knowledge across clients. A common design in PFL is the head-representation architecture, which combines a shared global representation with a local head on each client. Although effective, deploying this architecture in real-world systems remains challenging due to the presence of stragglers and the high communication cost. To address these issues, we propose FedSTEP, a unified framework that integrates asynchronous training with dynamic communication sparsification. Specifically, it adaptively adjusts each client's local training duration and communication sparsity based on staleness, enabling more efficient coordination between local adaptation and global representation. This design mitigates the impact of stragglers and ensures robust performance in heterogeneous environments. We provide a theoretical analysis of the convergence behavior and communication efficiency of FedSTEP under standard assumptions. Extensive experiments on five public datasets demonstrate that FedSTEP consistently outperforms existing methods. It achieves up to 4.65% higher accuracy, a 3.68× speedup in training, and a 1.91× reduction in communication cost. Gang Yan 0002, Jian Li 0008, Wan Du |
CIKM | 1 |
| 2025 | FedDiAL: Adaptive Federated Learning with Hierarchical Discriminative Network for Large Pre-trained ModelsabstractLarge pre-trained models have significantly advanced computer vision (CV) and natural language processing (NLP). However, their deployment is challenging in privacy-sensitive scenarios where data must remain decentralized. Federated Learning (FL) addresses this issue by enabling local model training without sharing raw data. Despite this advantage, integrating large models into FL introduces challenges such as training inefficiencies, label scarcity, and data heterogeneity. In this paper, we propose FedDiAL, a framework designed to effectively incorporate large pre-trained models into FL while addressing these challenges. At its core, FedDiAL, features the novel HDis-Net, which enhances the efficiency of large models in resource-constrained environments. We introduce a two-phase training strategy with probabilistic feature augmentation to improve feature discrimination. Additionally, we propose an adaptive pseudo-labeling method to generate high-confidence labels, mitigating label scarcity. To handle data heterogeneity, we develop a focused fine-tuning strategy that adapts HDis-Net to diverse client data distributions. Our theoretical analysis establishes the convergence of HDis-Net. Extensive experiments on four datasets, including Tiny ImageNet and AG News, demonstrate that FedDiAL outperforms state-of-the-art methods, achieving up to a 35.66% accuracy improvement in both CV and NLP tasks. Gang Yan 0002, Wan Du |
KDD (2) | 1 |
| 2024 | DePRL: Achieving Linear Convergence Speedup in Personalized Decentralized Learning with Shared RepresentationsabstractDecentralized learning has emerged as an alternative method to the popular parameter-server framework which suffers from high communication burden, single-point failure and scalability issues due to the need of a central server. However, most existing works focus on a single shared model for all workers regardless of the data heterogeneity problem, rendering the resulting model performing poorly on individual workers. In this work, we propose a novel personalized decentralized learning algorithm named DePRL via shared representations. Our algorithm relies on ideas from representation learning theory to learn a low-dimensional global representation collaboratively among all workers in a fully decentralized manner, as well as a user-specific low-dimensional local head leading to a personalized solution for each worker. We show that DePRL achieves, for the first time, a provable \textit{linear speedup for convergence} with general non-linear representations (i.e., the convergence rate is improved linearly with respect to the number of workers). Experimental results support our theoretical findings showing the superiority of our method in data heterogeneous environments. Guojun Xiong, Gang Yan 0002, Shiqiang Wang 0001, Jian Li 0008 |
AAAI | 2 |
| 2024 | FedRoLA: Robust Federated Learning Against Model Poisoning via Layer-based AggregationabstractFederated Learning (FL) is increasingly vulnerable to model poisoning attacks, where malicious clients degrade the global model's accuracy with manipulated updates. Unfortunately, most existing defenses struggle to handle the scenarios when multiple adversaries exist, and often rely on historical or validation data, rendering them ill-suited for the dynamic and diverse nature of real-world FL environments. Exacerbating these limitations is the fact that most existing defenses also fail to account for the distinctive contributions of Deep Neural Network (DNN) layers in detecting malicious activity, leading to the unnecessary rejection of benign updates. To bridge these gaps, we introduce FedRoLa, a cutting-edge similarity-based defense method optimized for FL. Specifically, FedRoLa leverages global model parameters and client updates independently, moving away from reliance on historical or validation data. It features a unique layer-based aggregation with dynamic layer selection, enhancing threat detection, and includes a dynamic probability method for balanced security and model performance. Through comprehensive evaluations using different DNN models and real-world datasets, FedRoLa demonstrates substantial improvements over the status quo approaches in global model accuracy, achieving up to 4% enhancement in terms of accuracy, reducing false positives to 6.4%, and securing an 92.8% true positive rate. Gang Yan 0002, Hao Wang 0022, Xu Yuan 0001, Jian Li 0008 |
KDD | 1 |
| 2024 | Straggler-Resilient Decentralized Learning via Adaptive Asynchronous UpdatesabstractWith the increasing demand for large-scale training of machine learning models, fully decentralized optimization methods have recently been advocated as alternatives to the popular parameter server framework. In this paradigm, each worker maintains a local estimate of the optimal parameter vector, and iteratively updates it by waiting and averaging all estimates obtained from its neighbors, and then corrects it on the basis of its local dataset. However, the synchronization phase is sensitive to stragglers. An efficient way to mitigate this effect is to consider asynchronous updates, where each worker computes stochastic gradients and communicates with other workers at its own pace. Unfortunately, fully asynchronous updates suffer from staleness of stragglers' parameters. To address these limitations, we propose a fully decentralized algorithm DSGD-AAU with adaptive asynchronous updates via adaptively determining the number of neighbor workers for each worker to communicate with. We show that DSGD-AAU achieves a linear speedup for convergence (i.e., convergence performance increases linearly with respect to the number of workers). Experimental results on a suite of datasets and deep neural network models are provided to verify our theoretical results. Guojun Xiong, Gang Yan 0002, Shiqiang Wang 0001, Jian Li 0008 |
MobiHoc | 2 |
| 2024 | Enhancing Model Poisoning Attacks to Byzantine-Robust Federated Learning via Critical Learning PeriodsabstractMost existing model poisoning attacks in federated learning (FL) control a set of malicious clients and share a fixed number of malicious gradients with the server in each FL training round, to achieve a desired tradeoff between the attack impact and the attack budget. In this paper, we show that such a tradeoff is not fundamental and an adaptive attack budget not only improves the impact of attack <?TeX $\mathcal {A}$?> Math 1 but also makes it more resilient to defenses. However, adaptively determining the number of malicious clients that share malicious gradients with the central server in each FL training round has been less investigated. This is due to the fact that most existing model poisoning attacks mainly focus on FL optimization itself to maximize the damage to the global model, and largely ignore the impact of the underlying deep neural networks that are used to train FL models. Inspired by recent findings on critical learning periods (CLP), where small gradient errors have irrecoverable impact on model accuracy, we advocate CLP augmented model poisoning attacks <?TeX $\mathcal {A}$?> Math 2 -CLP in this paper. <?TeX $\mathcal {A}$?> Math 3 -CLP merely augments an existing model poisoning attack <?TeX $\mathcal {A}$?> Math 4 with an adaptive attack budget scheme. Specifically, <?TeX $\mathcal {A}$?> Math 5 -CLP inspects the changes in federated gradient norms to identify CLP and adaptively adjusts the number of malicious clients that share their malicious gradients with the server in each round, leading to dramatically improved attack impact compared to <?TeX $\mathcal {A}$?> Math 6 by up to 6.85 ×, with a smaller attack budget. This in turn improves the resilience of <?TeX $\mathcal {A}$?> Math 7 by up to 2 ×. Since <?TeX $\mathcal {A}$?> Math 8 -CLP is orthogonal to the attack <?TeX $\mathcal {A}$?> Math 9 , it also crafts malicious gradients by solving a difficult optimization problem. To tackle this challenge and based on our understandings of <?TeX $\mathcal {A}$?> Math 10 -CLP, we further relax the inner attack subroutine <?TeX $\mathcal {A}$?> Math 11 in <?TeX $\mathcal {A}$?> Math 12 -CLP and design GraSP, a lightweight CLP augmented similarity-based attack. We show that GraSP not only is more flexible but also achieves an improved attack impact compared to the strongest of existing model poisoning attacks. Gang Yan 0002, Hao Wang 0022, Xu Yuan 0001, Jian Li 0008 |
RAID | 1 |
| 2023 | DeFL: Defending against Model Poisoning Attacks in Federated Learning via Critical Learning Periods AwarenessabstractFederated learning (FL) is known to be susceptible to model poisoning attacks in which malicious clients hamper the accuracy of the global model by sending manipulated model updates to the central server during the FL training process. Existing defenses mainly focus on Byzantine-robust FL aggregations, and largely ignore the impact of the underlying deep neural network (DNN) that is used to FL training. Inspired by recent findings on critical learning periods (CLP) in DNNs, where small gradient errors have irrecoverable impact on the final model accuracy, we propose a new defense, called a CLP-aware defense against poisoning of FL (DeFL). The key idea of DeFL is to measure fine-grained differences between DNN model updates via an easy-to-compute federated gradient norm vector (FGNV) metric. Using FGNV, DeFL simultaneously detects malicious clients and identifies CLP, which in turn is leveraged to guide the adaptive removal of detected malicious clients from aggregation. As a result, DeFL not only mitigates model poisoning attacks on the global model but also is robust to detection errors. Our extensive experiments on three benchmark datasets demonstrate that DeFL produces significant performance gain over conventional defenses against state-of-the-art model poisoning attacks. Gang Yan 0002, Hao Wang 0022, Xu Yuan 0001, Jian Li 0008 |
AAAI | 1 |
| 2023 | CriticalFL: A Critical Learning Periods Augmented Client Selection Framework for Efficient Federated LearningabstractFederated learning (FL) is a distributed optimization paradigm that learns from data samples distributed across a number of clients. Adaptive client selection that is cognizant of the training progress of clients has become a major trend to improve FL efficiency but not yet well-understood. Most existing FL methods such as FedAvg and its state-of-the-art variants implicitly assume that all learning phases during the FL training process are equally important. Unfortunately, this assumption has been revealed to be invalid due to recent findings on critical learning periods (CLP), in which small gradient errors may lead to an irrecoverable deficiency on final test accuracy. In this paper, we develop CriticalFL, a CLP augmented FL framework to reveal that adaptively augmenting exiting FL methods with CLP, the resultant performance is significantly improved when the client selection is guided by the discovered CLP. Experiments based on various machine learning models and datasets validate that the proposed CriticalFL framework consistently achieves an improved model accuracy while maintains better communication efficiency as compared to state-of-the-art methods, demonstrating a promising and easily adopted method for tackling the heterogeneity of FL training. Gang Yan 0002, Hao Wang 0022, Xu Yuan 0001, Jian Li 0008 |
KDD | 1 |
| 2023 | Reinforcement Learning for Dynamic Dimensioning of Cloud Caches: A Restless Bandit ApproachabstractWe study the dynamic cache dimensioning problem, where the objective is to decide how much storage to place in the cache to minimize the total costs with respect to the storage and content delivery latency. We formulate this problem as a Markov decision process, which turns out to be a restless multi-armed bandit problem and is provably hard to solve. For given dimensioning decisions, it is possible to develop solutions based on the celebrated Whittle index policy. However, Whittle index policy has not been studied for dynamic cache dimensioning, mainly because cache dimensioning needs to be repeatedly solved and jointly optimized with content caching. To overcome this difficulty, we propose a low-complexity fluid Whittle index policy, which jointly determines dimensioning and content caching. We show that this policy is asymptotically optimal. We further develop a lightweight reinforcement learning augmented algorithm dubbed fW-UCB when the content request and delivery rates are unavailable. fW-UCB is shown to achieve a sub-linear regret as it fully exploits the structure of the near-optimal fluid Whittle index policy and hence can be easily implemented. Extensive simulations using real traces support our theoretical results. Guojun Xiong, Shufan Wang, Gang Yan 0002, Jian Li 0008 |
IEEE/ACM Trans. Netw. | 3 |
| 2022 | Seizing Critical Learning Periods in Federated LearningabstractFederated learning (FL) is a popular technique to train machine learning (ML) models with decentralized data. Extensive works have studied the performance of the global model; however, it is still unclear how the training process affects the final test accuracy. Exacerbating this problem is the fact that FL executions differ significantly from traditional ML with heterogeneous data characteristics across clients, involving more hyperparameters. In this work, we show that the final test accuracy of FL is dramatically affected by the early phase of the training process, i.e., FL exhibits critical learning periods, in which small gradient errors can have irrecoverable impact on the final test accuracy. To further explain this phenomenon, we generalize the trace of the Fisher Information Matrix (FIM) to FL and define a new notation called FedFIM, a quantity reflecting the local curvature of each clients from the beginning of the training in FL. Our findings suggest that the initial learning phase plays a critical role in understanding the FL performance. This is in contrast to many existing works which generally do not connect the final accuracy of FL to the early phase training. Finally, seizing critical learning periods in FL is of independent interest and could be useful for other problems such as the choices of hyperparameters including but not limited to the number of client selected per round, batch size, so as to improve the performance of FL training and testing. Gang Yan 0002, Hao Wang 0022, Jian Li 0008 |
AAAI | 1 |
| 2022 | Reinforcement Learning for Dynamic Dimensioning of Cloud Caches: A Restless Bandit ApproachabstractWe study the dynamic cache dimensioning problem, where the objective is to decide how much storage to place in the cache to minimize the total costs with respect to the storage and content delivery latency. We formulate this problem as a Markov decision process, which turns out to be a restless multi-armed bandit problem and is provably hard to solve. For given dimensioning decisions, it is possible to develop solutions based on the celebrated Whittle index policy. However, Whittle index policy has not been studied for dynamic cache dimensioning, mainly because cache dimensioning needs to be repeatedly solved and jointly optimized with content caching. To overcome this difficulty, we propose a low-complexity fluid Whittle index policy, which jointly determines dimensioning and content caching. We show that this policy is asymptotically optimal. We further develop a lightweight reinforcement learning augmented algorithm dubbed fW-UCB when the content request and delivery rates are unavailable. fW-UCB is shown to achieve a sub-linear regret as it fully exploits the structure of the near-optimal fluid Whittle index policy and hence can be easily implemented. Extensive simulations using real traces support our theoretical results. Guojun Xiong, Shufan Wang, Gang Yan 0002, Jian Li 0008 |
INFOCOM | 3 |
| 2022 | Towards Latency Awareness for Content Delivery Network Caching
Gang Yan 0002, Jian Li 0008 |
USENIX ATC | 1 |
| 2021 | Learning from optimal caching for content deliveryabstractContent delivery networks (CDNs) distribute much of today's Internet traffic by caching and serving users' contents requested. A major goal of a CDN is to improve hit probabilities of its caches, thereby reducing WAN traffic and user-perceived latency. In this paper, we develop a new approach for caching in CDNs that learns from optimal caching for decision making. To attain this goal, we first propose HRO to compute the upper bound on optimal caching in an online manner, and then leverage HRO to inform future content admission and eviction. We call this new cache design LHR. We show that LHR is efficient since it includes a detection mechanism for model update, an auto-tuned threshold-based model for content admission with a simple eviction rule. We have implemented an LHR simulator as well as a prototype within an Apache Traffic Server and the Caffeine, respectively. Our experimental results using four production CDN traces show that LHR consistently outperforms state of the arts with an increase in hit probability of up to 9% and a reduction in WAN traffic of up to 15% compared to a typical production CDN cache. Our evaluation of the LHR prototype shows that it only imposes a moderate overhead and can be deployed on today's CDN servers. Gang Yan 0002, Jian Li 0008, Don Towsley |
CoNEXT | 1 |
| 2020 | RL-Bélády: A Unified Learning Framework for Content CachingabstractContent streaming is the dominant application in today's Internet, which is typically distributed via content delivery networks (CDNs). CDNs usually use caching as a means to reduce user access latency so as to enable faster content downloads. Typical analysis of caching systems either focuses on content admission, which decides whether to cache a content, or content eviction to decide which content to evict when the cache is full. This paper instead proposes a novel framework that can simultaneously learn both content admission and content eviction for caching in CDNs. To attain this goal, we first put forward a lightweight architecture for content next request time prediction. We then leverage reinforcement learning (RL) along with the prediction to learn the time-varying content popularities for content admission, and develop a simple threshold-based model for content eviction. We call this new algorithm RL-Bélády (RLB). In addition, we address several key challenges to design learning-based caching algorithms, including how to guarantee lightweight training and prediction with both content eviction and admission in consideration, limit memory overhead, reduce randomness and improve robustness in RL stochastic optimization. Our evaluation results using $3$ production CDN datasets show that RLB can consistently outperform state-of-the-art methods with dramatically reduced running time and modest overhead. Gang Yan 0002, Jian Li 0008 |
ACM Multimedia | 1 |