VLDB 2026 Research / reviewers in the wild / expert
Guangfeng Yan
dblp:268/0875
· DBLP profile ↗
10ranked-venue papers
6as first author
9since 2021 · last 2025
0000-0002-0936-9467ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 4 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Theory of computation · 2 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Layered Randomized Quantization for Communication-Efficient and Privacy-Preserving Distributed LearningabstractIn distributed learning systems, ensuring efficient communication and privacy protection are two significant challenges. Although several existing works have attempted to address these challenges simultaneously, they often overlook essential learning-oriented features such as dynamic gradient and communication characteristics. In this paper, we propose a communication-efficient and privacy-preserving distributed SGD algorithm. Our proposed algorithm employs a layered randomized quantizer (LRQ) to reduce communication overhead, which also ensures that quantization errors follow an exact Gaussian distribution, thus achieving client-level differential privacy. We analyze the trade-off between convergence error, communication, and privacy under non-IID data distributions. Besides, we modify the algorithm to be training-adaptive by adjusting the perround privacy budget allocation in response to i) dynamic gradient features and ii) real-time changing communication rounds. Both closed-form solutions are derived by solving the minimization problem of convergence error subject to the privacy budget constraint. Finally, we evaluate the effectiveness of our approach through extensive experiments on various datasets, including MNIST, CIFAR-10, and CIFAR-100, demonstrating its superiority in terms of communication cost, privacy protection, and model performance compared to state-of-the-art methods. Guangfeng Yan, Tan Li 0002, Kui Wu 0001, Linqi Song |
IEEE J. Sel. Areas Commun. | 1 |
| 2024 | Towards Quantum-Safe Distributed Learning via Homomorphic Encryption: Learning with GradientsabstractThis paper introduces a privacy-preserving distributed learning framework via private-key homomorphic encryption. Using randomness in the quantization of gradients, our encryption replaces the Gaussian error term of Learning With Errors (LWE) with quantized gradients, thus reducing the error expansion speed in conventional LWE-based homomorphic en-cryption. The proposed system allows a large number of learning participants to engage in distributed learning collaboratively over an honest-but-curious server, while ensuring the cryptographic security of participants' uploaded gradients. Guangfeng Yan, Shanxiang Lyu, Hanxu Hou, Zhiyong Zheng, Linqi Song |
ITW | 1 |
| 2024 | Truncated Non-Uniform Quantization for Distributed SGDabstractTo address the communication bottleneck challenge in distributed learning, our work introduces a novel two-stage quantization strategy designed to enhance the communication efficiency of distributed Stochastic Gradient Descent (SGD). The proposed method initially employs truncation to mitigate the impact of long-tail noise, followed by a non-uniform quantization of the post-truncation gradients based on their statistical characteristics. We provide a comprehensive convergence analysis of the quantized distributed SGD, establishing theoretical guarantees for its performance. Furthermore, by minimizing the convergence error, we derive optimal closed-form solutions for the truncation threshold and non-uniform quantization levels under given communication constraints. Both theoretical insights and extensive experimental evaluations demonstrate that our proposed algorithm outperforms existing quantization schemes, striking a superior balance between communication efficiency and convergence performance. Guangfeng Yan, Tan Li 0002, Yuanzhang Xiao, Hanxu Hou, Congduan Li, Linqi Song |
ITW | 1 |
| 2024 | Optimal selection of key parameters for homomorphic filtering based on information entropy
Zhantao Yang, Yangtenglong Li, Xuan Bai, Guangfeng Yan, Cong Fu 0015 |
Multim. Tools Appl. | 4 |
| 2023 | Adaptive Top- K in SGD for Communication-Efficient Distributed LearningabstractDistributed stochastic gradient descent (SGD) with gradient compression has become a popular communication-efficient solution for accelerating distributed learning. One commonly used method for gradient compression is Top-K sparsification, which sparsifies the gradients by a fixed degree during model training. However, there has been a lack of an adaptive approach to adjust the sparsification degree to maximize the potential of the model's performance or training speed. This paper proposes a novel adaptive Top-K in SGD framework that enables an adaptive degree of sparsification for each gradient descent step to optimize the convergence performance by balancing the tradeoff between communication cost and convergence error. Firstly, an upper bound of convergence error is derived for the adaptive sparsification scheme and the loss function. Secondly, an algorithm is designed to minimize the convergence error under the communication cost constraints. Finally, numerical results on the MNIST and CIFAR-10 datasets demonstrate that the proposed adaptive Top-K algorithm in SGD achieves a significantly better convergence rate compared to state-of-the-art methods, even after considering error compensation. Mengzhe Ruan, Guangfeng Yan, Yuanzhang Xiao, Linqi Song, Weitao Xu |
GLOBECOM | 2 |
| 2023 | Mitigating Model Poisoning Attacks on Distributed Learning with Heterogeneous DataabstractGradient-based distributed learning techniques have been essential for machine model training on distributed samples without collecting raw data. However, such learning systems are vulnerable to both internal failures and external attacks. In this work, we study the Byzantine robustness of distributed learning over heterogeneous data for classification tasks, where each worker only contains data samples belonging to some specific classes (aka., label skew) and a fraction of workers are corrupted by a Byzantine adversary to conduct attacks by sending malicious gradients. Existing defenses usually fail under such heterogeneous cases. To remedy this, we propose a gradient decomposition scheme called DeSGD to achieve more robust distributed model training. The key idea for mitigating the impact of data heterogeneity on the Byzantine robustness is to divide the full global gradient into individual gradients of each data class and conduct resilient aggregation in a class-wise manner. The proposed framework can easily integrate existing advanced defense methods and local momentum mechanism. Evaluation results on the Fashion-MNIST dataset with various strong attacks demonstrate the improved robustness of learning over distributed data in the presence of both label skew and attacks. Jian Xu 0016, Guangfeng Yan, Ziyan Zheng, Shao-Lun Huang |
ICMLA | 2 |
| 2022 | AC-SGD: Adaptively Compressed SGD for Communication-Efficient Distributed LearningabstractGradient compression (e.g., gradient quantization and gradient sparsification) is a core technique in reducing communication costs in distributed learning systems. The recent trend of gradient compression is to use a varying number of bits across iterations, however, relying on empirical observations or engineering heuristics without a systematic treatment and analysis. To the best of our knowledge, a general dynamic gradient compression that leverages both quantization and sparsification techniques is still far from understanding. This paper proposes a novel Adaptively-Compressed Stochastic Gradient Descent (AC-SGD) strategy to adjust the number of quantization bits and the sparsification size with respect to the norm of gradients, the communication budget, and the remaining number of iterations. In particular, we derive an upper bound, tight in some cases, of the convergence error for arbitrary dynamic compression strategy. Then we consider communication budget constraints and propose an optimization formulation - denoted as theAdaptive Compression Problem (ACP)- for minimizing the deep model’s convergence error under such constraints. By solving the ACP, we obtain an enhanced compression algorithm that significantly improves model accuracy under given communication budget constraints. Finally, through extensive experiments on computer vision and natural language processing tasks on MNIST, CIFAR-10, CIFAR-100 and AG-News datasets, respectively, we demonstrate that our compression scheme significantly outperforms the state-of-the-art gradient compression methods in terms of mitigating communication costs. Guangfeng Yan, Tan Li 0002, Shao-Lun Huang, Tian Lan 0001, Linqi Song |
IEEE J. Sel. Areas Commun. | 1 |
| 2022 | Communication-Efficient Federated Learning with Adaptive QuantizationabstractFederated learning (FL) has attracted tremendous attentions in recent years due to its privacy-preserving measures and great potential in some distributed but privacy-sensitive applications, such as finance and health. However, high communication overloads for transmitting high-dimensional networks and extra security masks remain a bottleneck of FL. This article proposes a communication-efficient FL framework with an Adaptive Quantized Gradient (AQG), which adaptively adjusts the quantization level based on a local gradient’s update to fully utilize the heterogeneity of local data distribution for reducing unnecessary transmissions. In addition, client dropout issues are taken into account and an Augmented AQG is developed, which could limit the dropout noise with an appropriate amplification mechanism for transmitted gradients. Theoretical analysis and experiment results show that the proposed AQG leads to 18% to 50% of additional transmission reduction as compared with existing popular methods, including Quantized Gradient Descent (QGD) and Lazily Aggregated Quantized (LAQ) gradient-based methods without deteriorating convergence properties. Experiments with heterogenous data distributions corroborate a more significant transmission reduction compared with independent identical data distributions. The proposed AQG is robust to a client dropping rate up to 90% empirically, and the Augmented AQG manages to further improve the FL system’s communication efficiency with the presence of moderate-scale client dropouts commonly seen in practical FL scenarios. Yuzhu Mao, Zihao Zhao 0001, Guangfeng Yan, Yang Liu 0165, Tian Lan 0001, Linqi Song, Wenbo Ding 0001 |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2021 | DQ-SGD: Dynamic Quantization in SGD for Communication-Efficient Distributed LearningabstractGradient quantization is an emerging technique in reducing communication costs in distributed learning. Existing gradient quantization algorithms often rely on engineering heuristics or empirical observations, lacking a systematic approach to dynamically quantize gradients. This paper addresses this issue by proposing a novel dynamically quantized SGD (DQ-SGD) framework, enabling us to dynamically adjust the quantization scheme for each gradient descent step by exploring the trade-off between communication cost and convergence error. We derive an upper bound, tight in some cases, of the convergence error for a restricted family of quantization schemes and loss functions. We design our DQSGD algorithm via minimizing the communication cost under the convergence error constraints. Finally, through extensive experiments on large-scale natural language processing and computer vision tasks on AG-News, CFAR-10, and CIFAR-100 datasets, we demonstrate that our quantization scheme achieves better tradeoffs between the communication cost and learning performance than other state-of-the-art gradient quantization methods. Guangfeng Yan, Shao-Lun Huang, Tian Lan 0001, Linqi Song |
MASS | 1 |
| 2020 | Unknown Intent Detection Using Gaussian Mixture Model with an Application to Zero-shot Intent ClassificationabstractUser intent classification plays a vital role in dialogue systems.Since user intent may frequently change over time in many realistic scenarios, unknown (new) intent detection has become an essential problem, where the study has just begun.This paper proposes Understanding user intent is crucial for developing conversational and dialogue systems.It is essential to accurately identify the intent behind a user utterance to better guide downstream decisions and policies.With the advent of conversational AI, dialogue systems are becoming central tools in many applications such as mobile apps, companion bots, virtual assistants and so on.Since user interests may change frequently over time, the AI agents may continuously see unknown (new) user intents.Manual annotation can hardly catch up with such rapid development, which motivates the problem * Equal contribution. Guangfeng Yan, Qimai Li, Han Liu 0008, Xiaotong Zhang 0003, Xiao-Ming Wu 0003, Albert Y. S. Lam |
ACL | 1 |