VLDB 2026 Research / reviewers in the wild / expert
Xiangrui Xu 0001
dblp:313/5210-1
· DBLP profile ↗
11ranked-venue papers
6as first author
11since 2021 · last 2025
0000-0003-3759-4706ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 7 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | From Risk to Resilience: Towards Assessing and Mitigating the Risk of Data Reconstruction Attacks in Federated Learning
Xiangrui Xu 0001, Zhize Li 0001, Yufei Han 0001, Bin Wang 0062, Jiqiang Liu, Wei Wang 0012 |
USENIX Security Symposium | 1 |
| 2025 | Enhancing Privacy in Distributed Intelligent Vehicles With Information Bottleneck TheoryabstractVertical federated learning (VFL) shows promise for enabling collaborative learning among Internet of Vehicle systems (IoVs) without requiring the sharing of private training data. However, existing work has exposed VFL’s vulnerability to privacy-stealing attacks, where an honest but curious server might reconstruct a client’s raw data from client-uploaded embeddings. In this work, we first elucidate the intrinsic mechanisms of privacy attacks from an information theory perspective, which provides a solid foundation for potential defensive strategies. Based on our findings, we introduce PriVFL, a defense mechanism based on information bottleneck theory. PriVFL is designed to safeguard the privacy of VFL-based IoVs by enabling shared embeddings to extract minimal information from input data, while preserving the information essential to target labels. Specifically, PriVFL restricts the information contained in embeddings by reducing the upper bound of mutual information between the raw samples and embeddings uploaded from local clients. Meanwhile, PriVFL ensures the effectiveness of the model by increasing the mutual information lower bound between embeddings and samples’ labels. Our evaluation includes 5 benchmark data sets and 4 different models. Experimental results demonstrate that PriVFL effectively mitigates privacy attacks while preserving the model’s effectiveness. These findings underscore that PriVFL can significantly enhance the privacy of VFL-based IoVs, thereby bolstering the development of practical IoV applications. Xiangrui Xu 0001, Pengrui Liu, Wei Wang 0012, Yongsheng Zhu, Chongzhen Zhang, Bin Wang 0062, Jian Shen 0001, Zhen Han 0001 |
IEEE Internet Things J. | 1 |
| 2025 | DMRP: Privacy-Preserving Deep Learning Model with Dynamic Masking and Random Permutation
Chongzhen Zhang, Zhiwang Hu, Xiangrui Xu 0001, Bin Wang 0062, Jian Shen 0001, Tao Li 0022, Baigen Cai, Wei Wang 0012 |
J. Inf. Secur. Appl. | 3 |
| 2025 | Finding the PISTE: Towards Understanding Privacy Leaks in Vertical Federated Learning SystemsabstractVertical Federated Learning (VFL) is a collaborative learning paradigm where participants share the same sample space while splitting the feature space. In VFL, local participants host their bottom models for feature extraction and collaboratively train a classifier by exchanging intermediate results with the server owning the labels. Both local training data and bottom models contain privacy-sensitive information and are considered the intellectual property of each participant, and thus should be protected by the design of VFL. Our study exposes the fundamental susceptibility of VFL systems to privacy leaks, which arise from the collaboration between the server and clients during both training and testing. Based on our findings, we proposePISTE, a model-agnostic framework of privacy stealing attacks against VFL. PISTE delivers three privacy inference attacks, i.e., model stealing, data reconstruction, and property inference attacks on five benchmark datasets and four different model architectures. We further discuss four potential countermeasures. Experimental results show that all of them cannot prevent all three privacy stealing attacks in PISTE. In summary, our study demonstrates the inherent yet rarely uncovered vulnerability of VFL on leaking data and model privacy. Xiangrui Xu 0001, Wei Wang 0012, Bin Wang 0062, Chao Li 0023, Zhen Han 0001, Yufei Han 0001 |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2025 | VFLMonitor: Defending One-Party Hijacking Attacks in Vertical Federated LearningabstractVertical Federated Learning (VFL) is susceptible to various one-party hijacking attacks, such as Replay and Generation attacks, where a single malicious client can manipulate the model to produce attacker-specified results, thereby compromising its reliability in real-world deployments. In this paper, we first uncover the underlying mechanisms of these attacks and observe that successful attacks induce significant discrepancies in the embedding-label associations across different clients. We establish a theoretical framework demonstrating how these discrepancies can serve as reliable indicators for detecting hijacking attempts. Building upon this insight, we propose VFLMonitor, a robust defense mechanism that leverages these embedding-label discrepancies to detect and mitigate hijacking attacks. Specifically, VFLMonitor identifies suspicious queries by analyzing differences in label estimations from multiple clients and applies a majority voting rule to correct or filter out these malicious queries. Moreover, VFLMonitor introduces a novel regularization strategy during training to reduce intra-class variance in embeddings, thereby enhancing their discriminative power and improving defense effectiveness. Extensive experi21 ments were conducted on 5 real-world datasets against 2 different attack types under 3 attack scenarios. The results demonstrate that VFLMonitor can effectively identify and exclude potential hijacked requests in all types of one-party hijacking attacks, while maintaining a meager false positive rate for legitimate queries. Xiangrui Xu 0001, Yufei Han 0001, Yongsheng Zhu, Zhen Han 0001, Guangquan Xu, Bin Wang 0062, Shouling Ji, Wei Wang 0012 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2025 | FedEditor: Efficient and Effective Federated Unlearning in Cooperative Intelligent Transportation SystemsabstractIn cooperative intelligent transportation systems (CITS), federated learning enables vehicles to train a global model without sharing private data. However, the lack of an unlearning mechanism to remove the influence of vehicle-specified data from the global model potentially violates data protection regulations regarding the right to be forgotten. While the existing federated unlearning (FU) methods exhibit promising unlearning effects, their practicality in CITS is hindered due to the time-consuming retraining steps required by other vehicles and the non-negligible performance sacrifice on the un-forgotten data. Therefore, achieving effective unlearning without extensive retraining, while minimizing performance degradation on the un-forgotten data remains a challenge. In this work, we propose FedEditor, an efficient and effective FU framework in CITS that addresses the above challenge by reconfiguring the global model’s representation space to remove critical classification-related knowledge from the unlearned data. Firstly, FedEditor enables vehicles to perform the unlearning process locally on the global model, eliminating the participation of other vehicles and improving efficiency. Secondly, FedEditor captures and aligns the representations of the unlearned data with those of the nearest incorrect class centroid derived from non-training data, ensuring effective unlearning while preserving the un-forgotten data’s knowledge relatively intact for achieving competitive model performance. Finally, FedEditor refines the global model’s output distributions using the vehicles’ remaining data and incorporates a drift-mitigating regularization term, minimizing the negative impact of unlearning operations on model performance. Experimental results show that FedEditor reduces the unlearning rate by up to 99.64% without time-consuming retraining, while limiting the predictive performance loss of the resulting global model to less than 3.88% across five models and seven datasets. Jiqiang Liu, Bin Wang 0051, Xiangrui Xu 0001, Tao Li 0022, Wei Wang 0012 |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2024 | Assessing Membership Leakages via Task-Aligned Divergent Shadow Data Sets in Vehicular Road CooperationabstractDeep classification models have been widely utilized in Vehicular Road Cooperation. However, previous work indicates that deep classification models are vulnerable to the privacy risks of Membership Inference Attacks (MIAs). Most existing work of MIAs is based on two different assumptions. One assumes adversary-own shadow datasets with aligned tasks and distributions as private datasets, while this assumption necessitates that the adversary knows the distributions of private datasets. The other assumes adversary-own shadow datasets with distinct tasks and distributions from private datasets, while this assumption requires that the adversary knows the classification boundaries between members and non-members of the private dataset. Hence, these two assumptions do not always hold in real-world scenarios. In this work, we systematically assess the impact of adversaryown shadow datasets with aligned tasks but distinct distributions from private datasets on MIAs. These realistic shadow datasets acknowledge adversaries limited insights into the data distribution of the private dataset and the decision boundary between members and non-members. We divide these practical shadow datasets hosted by adversaries into 4 types: Data Noise, Label Noise, Imbalanced Data, and Cross-domain Data. We conduct extensive experiments with 7 prevalent MIAs and 4 types of shadow datasets. Experimental results mainly reveal two-fold findings. First, MIAs still maintain effectiveness using the shadow dataset with the aligned task but distinct distributions from the private dataset. Second, different levels of data distribution disparities manifest varying MIAs’ performances under certain types of shadow datasets. Pengrui Liu, Wei Wang 0012, Xiangrui Xu 0001, Weiping Ding 0001 |
IEEE Internet Things J. | 3 |
| 2023 | CGIR: Conditional Generative Instance Reconstruction Attacks Against Federated LearningabstractData reconstruction attack has become an emerging privacy threat to Federal Learning (FL), inspiring a rethinking of FL's ability to protect privacy. While existing data reconstruction attacks have shown some effective performance, prior arts rely on different strong assumptions to guide the reconstruction process. In this work, we propose a novel Conditional Generative Instance Reconstruction Attack (CGIR attack) that drops all these assumptions. Specifically, we propose a batch label inference attack in non-IID FL scenarios, where multiple images can share the same labels. Based on the inferred labels, we conduct a “coarse-to-fine” image reconstruction process that provides a stable and effective data reconstruction. In addition, we equip the generator with a label condition restriction so that the contents and the labels of the reconstructed images are consistent. Our extensive evaluation results on two model architectures and five image datasets show that without the auxiliary assumptions, the CGIR attack outperforms the prior arts, even for complex datasets, deep models, and large batch sizes. Furthermore, we evaluate several existing defense methods. The experimental results suggest that pruning gradients can be used as a strategy to mitigate privacy risks in FL if a model tolerates a slight accuracy loss. Xiangrui Xu 0001, Pengrui Liu, Wei Wang 0012, Hongliang Ma, Bin Wang 0062, Zhen Han 0001, Yufei Han 0001 |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2022 | AdvCat: Domain-Agnostic Robustness Assessment for Cybersecurity-Critical Applications with Categorical InputsabstractMachine Learning-as-a-Service systems (MLaaS) have been largely developed for cybersecurity-critical applications, such as detecting network intrusions and fake news campaigns. Despite effectiveness, their robustness against adversarial attacks is one of the key trust concerns for MLaaS deployment. We are thus motivated to assess the adversarial robustness of the Machine Learning models residing at the core of these securitycritical applications with categorical inputs. Previous research efforts on accessing model robustness against manipulation of categorical inputs are specific to use cases and heavily depend on domain knowledge, or require white-box access to the target ML model. Such limitations prevent the robustness assessment from being as a domain-agnostic service provided to various real-world applications. We propose a provably optimal yet computationally highly efficient adversarial robustness assessment protocol for a wide band of ML-driven cybersecurity-critical applications. We demonstrate the use of the domain-agnostic robustness assessment method with substantial experimental study on fake news detection and intrusion detection problems. Helene Orsini, Hongyan Bao, Yujun Zhou 0002, Xiangrui Xu 0001, Yufei Han 0001, Longyang Yi, Wei Wang 0012, Xin Gao 0001, Xiangliang Zhang 0001 |
IEEE Big Data | 4 |
| 2022 | Threats, attacks and defenses to federated learning: issues, taxonomy and perspectivesabstractAbstract Empirical attacks on Federated Learning (FL) systems indicate that FL is fraught with numerous attack surfaces throughout the FL execution. These attacks can not only cause models to fail in specific tasks, but also infer private information. While previous surveys have identified the risks, listed the attack methods available in the literature or provided a basic taxonomy to classify them, they mainly focused on the risks in the training phase of FL. In this work, we survey the threats, attacks and defenses to FL throughout the whole process of FL in three phases, including Data and Behavior Auditing Phase , Training Phase and Predicting Phase . We further provide a comprehensive analysis of these threats, attacks and defenses, and summarize their issues and taxonomy. Our work considers security and privacy of FL based on the viewpoint of the execution process of FL. We highlight that establishing a trusted FL requires adequate measures to mitigate security and privacy threats at each phase. Finally, we discuss the limitations of current attacks and defense approaches and provide an outlook on promising future research directions in FL. Pengrui Liu, Xiangrui Xu 0001, Wei Wang 0012 |
Cybersecur. | 2 |
| 2021 | Conditional image generation with One-Vs-All classifier
Xiangrui Xu 0001, Cao Yuan |
Neurocomputing | 1 |