VLDB 2026 Research / reviewers in the wild / expert
Xue Li 0034
dblp:181/2710-34
· DBLP profile ↗
10ranked-venue papers
0as first author
10since 2021 · last 2026
0009-0002-6082-5955ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 5 · 5 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Dynamic federated semi-supervised learning with flexible unlabeled sample selection
Siguang Chen, Yanyan Xia, Xue Li 0034, Chuanxin Zhao |
Pattern Recognit. | 3 |
| 2026 | Image Style Transfer-Empowered Federated Domain GeneralizationabstractAs a distributed machine learning paradigm, federated learning enables collaborative training among multiple clients while preserving data privacy. However, in practical applications, it faces the challenge of domain shift caused by data heterogeneity, which limits the generalization performance of the global model on unseen target domains. To address this issue, this paper proposes an image style transfer-empowered federated domain generalization method. Specifically, the method first enriches the domain diversity of local data through image style transfer techniques. Meanwhile, we introduce a predictive consistency regularization term into the optimization objective, it ensures the model maintains stable outputs when processing both original samples and their restyled versions, thereby mitigating the overfitting to local data domains and facilitating the learning of domain-invariant features. Furthermore, a generalization capability-aware aggregation weight optimization strategy is developed. By leveraging an unlabeled public dataset on the server side and its restyled versions to simulate unseen target domains, the strategy evaluates the generalization performance of client models and dynamically adjusts aggregation weights accordingly, which enhances the contribution of clients with higher generalization capabilities to the global model. Finally, experiments validate the effectiveness of both the local regularization and aggregation weight optimization strategies. On the PACS and Office-Home datasets, the proposed method achieves higher average test accuracy compared to baseline methods, along with faster convergence speed. Moreover, we additionally conduct experiments on the Camelyon17 tumor classification dataset, which further verify the robustness and practical applicability of the proposed method in real-world scenarios. Qian Wang 0028, Xue Li 0034, Siguang Chen |
IEEE Trans. Image Process. | 3 |
| 2026 | Robust Federated Learning With Double DenoisingabstractFederated learning (FL), as a representative distributed learning paradigm, has achieved remarkable success. However, most existing FL studies assume that each client holds correctly labeled data, whereas in reality, noisy labels are ubiquitous on the client side. To mitigate the adverse impact of noisy labels on FL performance, we propose a robust FL method with double denoising. Specifically, we first perform primary denoising based on the idea of cross-prediction, where two global models are trained and mutually used on each client to identify and filter out mislabeled samples. Next, after fine-tuning the models, we develop secondary denoising by detecting and removing residual noisy samples through clustering based on loss values. In addition, we design a noise-tolerant local training strategy that dynamically assesses the influence of noisy data on local updates and applies differentiated update rules to prevent overfitting. Finally, experimental results on three benchmark datasets, including the real-world noisy dataset Clothing1M, demonstrate that our method effectively removes label noise, delivering improved performance of the global model. Xinglong Wei, Siguang Chen, Xue Li 0034, Song-Le Chen |
IEEE Trans. Mob. Comput. | 3 |
| 2025 | Debiased Device Sampling for Federated Edge Learning in Wireless NetworksabstractAs a privacy-preserved distributed machine learning paradigm, federated edge learning (FEL) was designed to absorb knowledge from user devices to construct intelligent services without transmitting raw data. However, this paradigm depends on the local training and model parameter transmission of user devices, therefore the computing power, storage capacity and network resources of the devices become the key factors to achieve energy well-budgeted and timely message transmission FEL. While in the wireless networks, those resources for devices are normally heterogeneous or limited. This paper aims to offer tangible solutions for optimal convergence and Quality of Service (QoS) assurance of FEL in wireless networks. First, we define a mathematical model for energy-efficient message transmission of FEL and formulate an optimization problem involving device sampling and resource allocation to attain optimal training convergence within energy and time constraints. Second, we theoretically analyze the impact of limited resources on sampling strategies and training convergence, thus simplifying the optimization problem for solvability. Third, we introduce an iterative heuristic algorithm that utilizes available resources to reduce client sampling bias. Extensive experiments show that our method can effectively obtain the debiased sampling strategy, and outperforms similar methods by minimizing device disconnection due to energy use and enhancing model convergence and performance. Siguang Chen, Yanhang Shi, Xue Li 0034 |
IEEE Trans. Mob. Comput. | 4 |
| 2025 | Digital Twin-Empowered Federated Incremental Learning for Non-IID Privacy DataabstractFederated learning (FL) has emerged as a compelling distributed learning paradigm without sharing local original data. However, with ubiquitous non-independent and identically distributed (non-IID) privacy data, the FL suffers from severe performance loss and the privacy leakage by inference attacks. Existing solutions lack a cohesive framework with theoretical support, and their performance optimization and privacy protection are inter-inhibitive or high-cost. In this paper, we propose a digital twin (DT)-empowered federated incremental learning method to tackle the above challenges. First, we construct a DT-empowered federated incremental learning model to achieve cooperative awareness of performance and privacy-preservation. Second, a diffusion model-based selective data synthesis method is designed to provide auxiliary data for FL, it can avoid unnecessary overhead while ensuring the quality of synthetic samples under non-IID. Besides, it alleviates the negative impact of non-IID by allocating a class-balanced sub-dataset to each DT with IID setting. Third, we develop a DT-empowered alternating incremental learning method initiatively, under the premise of ensuring the confidentiality of original dataset, it can achieve efficient FL performance under non-IID with a small amount of synthetic samples. Furthermore, in order to estimate the contribution of each local model accurately, we investigate a comentropy-based federated aggregation strategy, which can obtain a superior global model. By sufficient theoretical analysis, we prove that the proposed methodology can achieve consistent enhancement of performance and privacy-preservation. Simultaneously, the experiments demonstrate that our methodology has efficient privacy-preserving property, it also outperforms other benchmarks on the accuracy and stability of the global model, especially in highly heterogeneous scenarios. Qian Wang 0028, Siguang Chen, Meng Wu 0003, Xue Li 0034 |
IEEE Trans. Mob. Comput. | 4 |
| 2025 | FedUP: Federated Unlearning With PrototypesabstractAs an extension of machine unlearning in distributed scenarios, federated unlearning gains significant attention. However, federated unlearning remains challenging, as many studies require additional resources, such as auxiliary dataset or storage, to achieve high-quality models. These requirements incur extra costs and are often difficult to satisfy in practical applications. To address these issues, we propose a flexible client-level federated unlearning algorithm with prototypes, called FedUP. Specifically, our algorithm consists of two components: prototype-based unlearning and model recovering. First, we design a prototype-based unlearning strategy that uses prototypes of the erased client to guide the unlearning process, and maximizes the prototype loss between the remaining and erased clients to unlearn the information. It does not rely on historical storage updates or additional standard datasets, making the unlearning process more streamlined. To mitigate performance degradation from the unlearning process, we develop a brief model recovering approach guided by global prototypes to swiftly and efficiently restore models' accuracy on the remaining datasets. Unlike other unlearning algorithms, our approach exchanges prototypes instead of model parameters, significantly reducing communication overhead. Finally, we empirically evaluate the proposed algorithm from multiple perspectives on two datasets, demonstrating that our algorithm can achieve high-quality unlearned models with minimal communication cost. Yuhong Huang, Xue Li 0034, Song-Le Chen, Siguang Chen |
IEEE Trans. Sustain. Comput. | 2 |
| 2024 | Compressed-Sensing-Based Practical and Efficient Privacy-Preserving Federated LearningabstractFederated learning (FL) is a popular distributed learning framework that is proposed to address privacy concerns in traditional machine learning. However, recent research has highlighted an issue where model or gradient updates can be exploited to infer sensitive information from the training data, resulting in severe privacy leakage. Existing defenses against gradient leakage attacks often suffer from high computation overhead or compromised model performance. In addition, most defense methods lack sufficient protection for labels. To overcome the above shortcomings, we develop a compressed sensing (CS)-based practical and efficient privacy-preserving FL scheme. In order to provide simultaneous protection for both original data and labels, we propose a CS-based gradient perturbation method, which eliminates the information in the gradient that is commonly exploited by attackers to extract labels, and increases the discrepancy between the perturbed and original gradients. Meanwhile, double aggregation is adopted together to ensure individual gradients are not easily disclosed by attackers. We also design a novel gradient reconstruction method that adaptively estimates the true gradient sparsity used for decompressing, thereby improving the model performance in practical scenarios. Furthermore, our CS-based gradient compression reduces communication overhead and requires low computation overhead as it only involves fast matrix multiplication. Extensive experiment results demonstrate the strong privacy protection effects of our proposed scheme compared to other approaches across various settings, with advantages in terms of communication overhead, computation overhead, and model accuracy. Siguang Chen, Yifeng Miao, Xue Li 0034, Chuanxin Zhao |
IEEE Internet Things J. | 3 |
| 2024 | Prototypes Contrastive Learning Empowered Intelligent Diagnosis for Skin LesionabstractFederated learning (FL) has been widely adopted for intelligent skin lesion diagnosis due to enabling collaborative model training across distributed data sources while preserving data privacy. However, traditional centralized learning paradigm, which collects data from distributed institutions to train neural models centrally, has the risk of privacy leakage. FL enables collaborative learning on training an artificial intelligence (AI) model among scattered medical institutions (MIs) without directly sharing raw skin lesion data, which is a promising solution. However, FL generally suffers from performance deterioration due to heterogeneous skin lesion data sets, and the differences in computing power and communication conditions among different MIs can impact the training efficiency. In this article, we present a prototypes contrastive learning (PCL) empowered intelligent diagnosis mechanism for skin lesion. Our mechanism designs a PCL-based local training to overcome data heterogeneity by designing two special losses, where prototypes-based contrastive loss can increase the interclass variance and reduce the intraclass variance of various skin lesions, and prototypes-based consistent regularization loss can prevent the shift of local updating. Meanwhile, a hierarchical FL aggregation mechanism is proposed, where asynchronous and synchronous aggregations are combined to ensure the training efficiency in case of large differences in computing power and communication conditions. Finally, simulation results show that our developed method can effectively mitigate the adverse impact of heterogeneity skin lesion data sets and provide efficient FL training. Congying Duan, Siguang Chen, Xue Li 0034 |
IEEE Internet Things J. | 3 |
| 2023 | Skin Lesion Intelligent Diagnosis in Edge Computing Networks: An FCL ApproachabstractIn recent years, automatic skin lesion diagnosis methods based on artificial intelligence have achieved great success. However, the lack of labeled data, visual similarity between skin diseases, and restriction on private data sharing remain the major challenges in skin lesion diagnosis. In this article, first, we propose a federated contrastive learning framework to break down data silos and enhance the generalizability of diagnostic model to unseen data. Subsequently, by combining data features from different participated nodes, the proposed framework can improve the performance of contrastive training. To extract discriminative features during on-device training, we propose a contrastive learning based intelligent skin lesion diagnosis scheme in edge computing networks. Specifically, a contrastive learning based dual encoder network is designed to overcome training sample scarcity by fully leveraging unlabeled samples for performance improvement. Meanwhile, we devise a maximum mean discrepancy based supervised contrastive loss function, which can efficiently explore complex intra-class and inter-class variances of samples. Finally, the diagnosis simulations demonstrate that compared with existing methods, our proposed scheme can achieve superior accuracy in both on-device training and distributed training scenarios. Yanhang Shi, Xue Li 0034, Siguang Chen |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2023 | Metadata and Image Features Co-Aware Personalized Federated Learning for Smart HealthcareabstractRecently, artificial intelligence has been widely used in intelligent disease diagnosis and has achieved great success. However, most of the works mainly rely on the extraction of image features but ignore the use of clinical text information of patients, which may limit the diagnosis accuracy fundamentally. In this paper, we propose a metadata and image features co-aware personalized federated learning scheme for smart healthcare. Specifically, we construct an intelligent diagnosis model, by which users can obtain fast and accurate diagnosis services. Meanwhile, a personalized federated learning scheme is designed to utilize the knowledge learned from other edge nodes with larger contributions and customize high-quality personalized classification models for each edge node. Subsequently, a Naïve Bayes classifier is devised for classifying patient metadata. And then the image and metadata diagnosis results are jointly aggregated by different weights to improve the accuracy of intelligent diagnosis. Finally, the simulation results illustrate that, compared with the existing methods, our proposed algorithm achieves better classification accuracy, reaching about 97.16% on PAD-UFES-20 dataset. Tong Jin 0005, Shujia Pan, Xue Li 0034, Siguang Chen |
IEEE J. Biomed. Health Informatics | 3 |