EDBT 2026 Demo / reviewers in the wild / expert
Wenjie Li 0008
dblp:33/3999-8
· DBLP profile ↗
14ranked-venue papers
5as first author
12since 2021 · last 2026
0000-0003-2717-9031ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 5 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 2 since 2021Computer networks · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Oblivionis: A Lightweight Learning and Unlearning Framework for Federated Large Language ModelsabstractLarge Language Models (LLMs) increasingly leverage Federated Learning (FL) to utilize private, task-specific datasets for fine-tuning while preserving data privacy. However, while federated LLM frameworks effectively enable collaborative training without raw data sharing, they critically lack built-in mechanisms for regulatory compliance like GDPR’s right to be forgotten. Integrating private data heightens concerns over data quality and long-term governance, yet existing distributed training frameworks offer no principled way to selectively remove specific client contributions post-training. Due to distributed data silos, stringent privacy constraints, and the intricacies of interdependent model aggregation, federated LLM unlearning is significantly more complex than centralized LLM unlearning. To address this gap, we introduce Oblivionis, a lightweight learning and unlearning framework that enables clients to selectively remove specific private data during federated LLM training, enhancing trustworthiness and regulatory compliance. By unifying FL and unlearning as a dual optimization objective, we incorporate 6 FL and 5 unlearning algorithms for comprehensive evaluation and comparative analysis, establishing a robust pipeline for federated LLM unlearning. Extensive experiments demonstrate that Oblivionis outperforms local training, achieving a robust balance between forgetting efficacy and model utility, with cross-algorithm comparisons providing clear directions for future LLM development. Fuyao Zhang, Xinyu Yan 0003, Tiantong Wu, Wenjie Li 0008, Yang Cao 0011, Longtao Huang, Wei Yang Bryan Lim, Qiang Yang 0001 |
AAAI | 4 |
| 2026 | Vertical Semi-Federated Learning for Efficient Online AdvertisingabstractTraditional vertical federated learning schema suffers from two main issues: 1) restricted applicable scope to overlapped samples and 2) high system challenge of real-time federated serving, which limits its application to advertising systems. To this end, we advocate a new practical learning setting, Semi-VFL (Vertical Semi-Federated Learning), for real-world industrial applications, where the learned model retains sufficient advantages of federated learning while supporting independent local serving. To achieve this goal, we propose the carefully designed Joint Privileged Learning framework (JPL) to i) alleviate the absence of the passive party's feature with federated equivalence imitation and ii) adapt to the heterogeneous full sample space with cross-branch rank alignment. Extensive experiments conducted on real-world advertising datasets validate the effectiveness of our method over baseline methods. Wenjie Li 0008, Shutao Xia, Jiangke Fan |
WWW | 1 |
| 2026 | NeuDFL: Efficient Neuron-Based Defense Against Label Flipping Attacks on Non-IID DataabstractFederated Learning (FL) enables collaborative model training while preserving data privacy, making it particularly attractive for large-scale Internet of Things (IoT) systems. However, in practical deployments, data collected by distributed clients are often non-independent and identically distributed (Non-IID), which amplifies the vulnerability of FL to poisoning attacks. Among them, Label-Flipping Attacks (LFA) are especially stealthy, as they can induce targeted misclassification without noticeably affecting overall accuracy, posing serious risks to safety-critical IoT applications. In this paper, we propose NeuDFL, a lightweight and robust defense framework against LFA under Non-IID settings. Unlike many existing defenses that primarily rely on gradient analysis over the full model update space or auxiliary clean datasets, NeuDFL exploits lightweight class-wise parameter statistics extracted from the final fully connected layer. By leveraging the cumulative and task-aligned nature of model parameters, NeuDFL enables reliable identification of attacked classes and filters malicious clients via an adaptive statistical threshold, improving robustness to data heterogeneity while incurring low computational overhead. Extensive experiments on multiple datasets demonstrate that NeuDFL offers an effective and efficient defense against label-flipping attacks, providing a robust solution for federated learning in complex real-world environments. Kai Fan 0001, Huixuan Wang, Wenjie Li 0008, Hui Li 0006, Kuan Zhang 0001, Yintang Yang, Lianhai Wang |
IEEE Internet Things J. | 3 |
| 2026 | DeepCGC: Unveiling the Deep Clustering Mechanism of Fast Graph CondensationabstractGraph condensation (GC) improves the efficiency of GNN training by condensing a large-scale graph into a compact synthetic graph. However, existing GC methods suffer from time-consuming optimization processes, and the underlying mechanisms driving their effectiveness remain unexplored. In this paper, we provide novel insights into the optimization strategies of GC, demonstrating that various methods ultimately converge to the class-level feature matching between the original and condensed graphs. Building on this understanding, we further refine the unified class-to-class matching paradigm into a fine-grained class-to-node paradigm, unveiling that the core mechanism of GC is a class-wise clustering problem in the latent space. Accordingly, we propose Deep Clustering-based Graph Condensation (DeepCGC), an efficient GC framework that integrates a clustering-based optimization objective with an invertible relay model. Extensive experiments show that DeepCGC achieves state-of-the-art efficiency and accuracy. Notably, it condenses the million-scale Ogbn-products graph in around 40 seconds—a$10^{2} \times$to$10^{4} \times$speedup over existing methods—while boosting accuracy by up to 4.6%. The code is available athttps://github.com/XYGaoG/DeepCGC. Xinyi Gao 0001, Wenjie Li 0008, Tong Chen 0005, Xiangyu Zhao 0001, Nguyen Quoc Viet Hung, Hongzhi Yin |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2025 | Secure and Scalable Face Retrieval via Cancelable Product Quantization
Haomiao Tang, Wenjie Li 0008, Yixiang Qiu, Genping Wang, Shutao Xia |
PRCV (6) | 2 |
| 2025 | Enhancing Security and Privacy in Federated Learning Using Low-Dimensional Update Representation and Proximity-Based DefenseabstractFederated Learning (FL) is a promising privacy-preserving machine learning paradigm that allows data owners to collaboratively train models while keeping their data localized. Despite its potential, FL faces challenges related to the trustworthiness of both clients and servers, particularly against curious or malicious adversaries. In this paper, we introduce a novel framework namedFederatedLearning with Low-DimensionalUpdateRepresentation andProximity-Based defense (FLURP), designed to address privacy preservation and resistance to Byzantine attacks in distributed learning environments. FLURP employs$\mathsf {LinfSample}$method, enabling clients to compute the$l_{\infty }$norm across sliding windows of updates, resulting in a Low-Dimensional Update Representation (LUR). Calculating the shared distance matrix among LURs, rather than updates, significantly reduces the overhead of Secure Multi-Party Computation (SMPC) by three orders of magnitude while effectively distinguishing between benign and poisoned updates. Additionally, FLURP integrates a privacy-preserving proximity-based defense mechanism utilizing optimized SMPC protocols to minimize communication rounds. Our experiments demonstrate FLURP's effectiveness in countering Byzantine adversaries with low communication and runtime overhead. FLURP offers a scalable framework for secure and reliable FL in distributed environments, facilitating its application in scenarios requiring robust data management and security. Wenjie Li 0008, Kai Fan 0001, Hui Li 0006, Wei Yang Bryan Lim, Qiang Yang 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2024 | ReFer: Retrieval-Enhanced Vertical Federated Recommendation for Full Set User Benefit
Wenjie Li 0008, Zhongren Wang 0003, Jinpeng Wang 0002, Shutao Xia, Jile Zhu, Jiangke Fan |
SIGIR | 1 |
| 2024 | FLSG: A Novel Defense Strategy Against Inference Attacks in Vertical Federated LearningabstractAs a new machine learning (ML) paradigm, federated learning (FL) empowers different participants to jointly train a more effective model than traditional ML. Unlike horizontal FL (HFL), which expands the sample space by aggregating local models, vertical FL (VFL) is suitable for scenarios where the sample ID between participants is the same but differs in sample characteristics. For a long time, VFL has been considered safe due to no data exchange and heterogeneity between parties. However, the recently proposed label inference attacks pose a significant security threat to VFL. Specifically, by adding randomly initialized layers to the top of local models, the passive label inference attack can infer tens of thousands of local participants’ private data with only 40 auxiliary labels. Since the attack is entirely local, using privacy protection technologies such as differential privacy cannot effectively defend against these attacks. Therefore, we propose a new privacy protection scheme called FL similar gradients (FLSGs) to defend against this attack. Unlike differential privacy, the FLSG scheme randomly generates gradients of a Gaussian distribution similar in dimension to the original gradients and calculates their cosine distance. If the distance is less than a certain threshold, the gradients are used instead of the original gradients to pass to the local participants. We conducted extensive evaluations on six real-world data sets, and the results show that FLSG provides a better defensive effect at lower computational overhead than other known methods when defending the passive label inference attack. Kai Fan 0001, Jingtao Hong, Wenjie Li 0008, Xingwen Zhao, Hui Li 0006, Yintang Yang |
IEEE Internet Things J. | 3 |
| 2024 | PBFL: Privacy-Preserving and Byzantine-Robust Federated-Learning-Empowered Industry 4.0abstractIn Industry 4.0, artificial intelligence (AI) has been successfully applied in scenarios, such as fault prediction, traffic analysis, and production decision making. However, due to the sensitivity and security of data, privacy regulations prohibit the transfer and exchange of industrial data between entities, resulting in training data being fragmented into data silos that limit the accuracy of AI models. FL can effectively break the data silo effect, but naive federated learning (FL) (FedAvg) is vulnerable to inference attacks from aggregators and Byzantine attacks from participants. To address these issues, we propose a privacy-preserving and Byzantine-robust federated learning scheme (PBFL) for Industry 4.0. Under the setting of an benign-majority participants, PBFL can always identify benign direction and magnitude of updates. Extensive experiments demonstrate that PBFL is more robust than state-of-the-art schemes, even with extreme proportion (49%) of malicious participants. Moreover, PBFL contains a series of well-optimized 2-party computation (2PC) protocols, causing it reduces total runtime of the unoptimized implementation by around$3 \times \sim 4 \times $and$9 \times \sim 10 \times $for 32-bit and 64-bit circuits, respectively. Wenjie Li 0008, Kai Fan 0001, Kan Yang 0001, Yintang Yang, Hui Li 0006 |
IEEE Internet Things J. | 1 |
| 2024 | A Game-theoretic Framework for Privacy-preserving Federated LearningabstractIn federated learning, benign participants aim to optimize a global model collaboratively. However, the risk of privacy leakage cannot be ignored in the presence of semi-honest adversaries. Existing research has focused either on designing protection mechanisms or on inventing attacking mechanisms. While the battle between defenders and attackers seems never-ending, we are concerned with one critical question: Is it possible to prevent potential attacks in advance? To address this, we propose the first game-theoretic framework that considers both FL defenders and attackers in terms of their respective payoffs, which include computational costs, FL model utilities, and privacy leakage risks. We name this game the federated learning privacy game (FLPG), in which neither defenders nor attackers are aware of all participants’ payoffs. To handle the incomplete information inherent in this situation, we propose associating the FLPG with an oracle that has two primary responsibilities. First, the oracle provides lower and upper bounds of the payoffs for the players. Second, the oracle acts as a correlation device, privately providing suggested actions to each player. With this novel framework, we analyze the optimal strategies of defenders and attackers. Furthermore, we derive and demonstrate conditions under which the attacker, as a rational decision-maker, should always follow the oracle’s suggestion not to attack . Xiaojin Zhang 0002, Lixin Fan, Wenjie Li 0008, Kai Chen 0005, Qiang Yang 0001 |
ACM Trans. Intell. Syst. Technol. | 4 |
| 2022 | Deep Dirichlet process mixture modelsabstractIn this paper we propose the deep Dirichlet process mixture (DDPM) model, which is an unsupervised method that simultaneously performs clustering and feature learning. The traditional Dirichlet process mixture model can infer the number of mixture components, but its flexibility is restricted since the clustering is performed in the raw feature space. Our method alleviates this limitation by using the flow-based deep neural network to learn more expressive features. DDPM unifies Dirichlet processes and the flow-based model with Monte Carlo expectation-maximization, and uses Gibbs sampling to sample from the posterior. This combination allows our method to exploit the mutually beneficial relation between clustering and feature learning. The effectiveness of DDPM is demonstrated by thorough experiments in various synthetic and real-world datasets. Naiqi Li, Wenjie Li 0008, Yong Jiang 0001, Shutao Xia |
UAI | 2 |
| 2021 | H-GPR: A Hybrid Strategy for Large-Scale Gaussian Process RegressionabstractWith the massive volume of data emerging from both scientific and industrial domains, it has become a desideratum to improve the scalability of Gaussian process regression (GPR). There are two major approaches to assuage its $\mathcal{O}\left( {{n^3}} \right)$ training complexity: the aggregation based methods and the sparse approximation methods. This paper proposes a hybrid strategy called H-GPR to combine these two well-established approaches. We show that it is possible to improve the performance of aggregation based methods by removing some data points that severely violate its underlying assumption, and then this information loss can be recovered by a set of inducing points generated by the sparse approximation methods. A novel metric called conditional independent score is proposed, which measures to what extent the assumption made by the aggregation based methods is satisfied. A heuristic rule is developed to adjust the relative size of the local experts and the inducing subset, so that their distinction can be better reflected. Thorough experiments on synthetic and realistic datasets were performed, demonstrating that the proposed method can improve both the predictive means and variances. Naiqi Li, Yinghua Gao, Wenjie Li 0008, Yong Jiang 0001, Shutao Xia |
ICASSP | 3 |
| 2020 | SDCN: Sparsity and Diversity Driven Correlation Networks for Traffic Demand ForecastingabstractTraffic demand forecasting is essential to intelligent transportation systems and is widely used to support urban planning, traffic management and vehicle dispatching. One challenge of this problem is to model the complex spatial-temporal correlation. Although both factors have been studied, many of the existing works have strong limitations. They rely too heavily on the locality assumption (i.e., the local area is more relevant than the remote area) and only use a distance-based correlation measurement. However, the spatial correlation is also global (i.e., areas far away may also be relevant) and sparse. And it's insufficient to measure the spatial correlation using only the distance measurement. In this paper, a sparsity and diversity driven correlation network is proposed to tackle these issues. Firstly, Multiple sparse correlation graphs are carefully generated to encode sparsity and diversity. Then a newly designed hybrid graph filtering module (HGFM) leverages them to learn a more expressive node representation. Finally, the HGFM-based recurrent filtering module (RFM) is introduced to handle the spatial-temporal correlation. Extensive experiments conducted on real-world datasets demonstrate the competitiveness of our model while showing the significance of sparsity and diversity. Wenjie Li 0008, Xue Yang 0003, Xiaohu Tang 0004, Shutao Xia |
IJCNN | 1 |
| 2020 | Stochastic Deep Gaussian Processes over GraphsabstractIn this paper we propose Stochastic Deep Gaussian Processes over Graphs (DGPG), which are deep structure models that learn the mappings between input and output signals in graph domains. The approximate posterior distributions of the latent variables are derived with variational inference, and the evidence lower bound is evaluated and optimized by the proposed recursive sampling scheme. The Bayesian non-parametric natural of our model allows it to resist overfitting, while the expressive deep structure grants it the potential to learn complex relations. Extensive experiments demonstrate that our method achieves superior performances in both small size (< 50) and large size (> 35,000) datasets. We show that DGPG outperforms another Gaussian-based approach, and is competitive to a state-of-the-art method in the challenging task of traffic flow prediction. Our model is also capable of capturing uncertainties in a mathematical principled way and automatically discovering which vertices and features are relevant to the prediction. Naiqi Li, Wenjie Li 0008, Jifeng Sun, Yinghua Gao, Yong Jiang 0001, Shutao Xia |
NeurIPS | 2 |