Xinjian Luo

dblp:146/1133 · DBLP profile ↗
← Back
14ranked-venue papers
9as first author
10since 2021 · last 2025
0000-0001-9671-5104ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 5 · 3 first-author · 5 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-authorTheory of computation · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 Prompt Inference Attack on Distributed Large Language Model Inference Frameworks
Xinjian Luo, Ting Yu 0001, Xiaokui Xiao
CCS1
2025 Passive Inference Attacks on Split Learning via Adversarial Regularization
Xiaochen Zhu 0003, Xinjian Luo, Yuncheng Wu, Yangfan Jiang 0001, Xiaokui Xiao, Beng Chin Ooi
NDSS2
2025 Exploiting Defenses against GAN-Based Feature Inference Attacks in Federated Learning
abstract
Federated Learning (FL) is a decentralized model training framework that aims to merge isolated data islands while maintaining data privacy. However, recent studies have revealed that Generative Adversarial Network (GAN)-based attacks can be employed in FL to learn the distribution of private datasets and reconstruct recognizable images. In this article, we exploit defenses against GAN-based attacks in FL and propose a framework, Anti-GAN, to prevent attackers from learning the real distribution of the victim’s data. The core idea of Anti-GAN is to manipulate the visual features of private training images to make them indistinguishable to human eyes even restored by attackers. Specifically, Anti-GAN projects the private dataset onto a GAN’s generator and combines the generated fake images with the actual images to create the training dataset, which is then used for federated model training. The experimental results demonstrate that Anti-GAN is effective in preventing attackers from learning the distribution of private images while causing minimal harm to the accuracy of the federated model.
Xinjian Luo, Xianglong Zhang
ACM Trans. Knowl. Discov. Data1
2024 Protecting Label Distribution in Cross-Silo Federated Learning
abstract
Federated learning (FL) is a popular distributed machine learning (ML) framework in which multiple parties share their model parameters instead of the raw training datasets to construct a global model in a privacy-preserving manner. However, existing FL solutions mainly focus on protecting the privacy of individual training records by incorporating differential privacy (DP), while overlooking the protection of the distribution information of training datasets, despite the fact that data distribution is also regarded as highly sensitive in high-stakes applications.In this paper, we propose the first privacy-preserving stochastic gradient descent (SGD) algorithm for protecting label distribution in FL. To establish a formal privacy guarantee, we formalize a privacy notion, dubbed (m,γ,ξ)-label distributional privacy, to quantify label distributional privacy leakage. Subsequently, we design the label distribution perturbation mechanism (LDPM) that carefully incorporates randomness into the SGD algorithm to achieve (m,γ,ξ)-label distributional privacy for all one-vs-all classification models. LDPM is easy to implement and provides non-trivial privacy guarantees, making it a suitable drop-in replacement for existing FL local model training algorithms. Notably, we demonstrate that LDPM also ensures DP, indicating that LDPM offers both individual and label distributional privacy guarantees. Extensive experiments on six benchmark datasets validate the effectiveness of LDPM.
Yangfan Jiang 0001, Xinjian Luo, Yuncheng Wu, Xiaokui Xiao, Beng Chin Ooi
SP2
2024 Calibrating Noise for Group Privacy in Subsampled Mechanisms
abstract
Given a group size m and a sensitive dataset D , group privacy (GP) releases information about D (e.g., weights of a neural network trained on D) with the guarantee that the adversary cannot infer with high confidence whether the underlying data is D or a neighboring dataset D ′ that differs from D by m records. GP generalizes the well-established notion of differential privacy (DP) for protecting individuals' privacy; in particular, when m = 1, GP reduces to DP. Compared to DP, GP is capable of protecting the sensitive aggregate information of a group of up to m individuals, e.g., the average annual income among members of a yacht club. Despite its longstanding presence in the research literature and its promising applications, GP is often treated as an afterthought, with most approaches first developing a differential privacy (DP) mechanism and then using a generic conversion to adapt it for GP, treating the DP solution as a black box. As we point out in the paper, this methodology is suboptimal when the underlying DP solution involves subsampling, e.g., in the classic DP-SGD method for training deep learning models. In this case, the DP-to-GP conversion is overly pessimistic in its analysis, leading to high error and low utility in the published results under GP. Motivated by this, we propose a novel analysis framework that provides tight privacy accounting for subsampled GP mechanisms. Instead of converting a black-box DP mechanism to GP, our solution carefully analyzes and utilizes the inherent randomness in subsampled mechanisms, leading to a substantially improved bound on the privacy loss with respect to GP. The proposed solution applies to a wide variety of foundational mechanisms with subsampling. Extensive experiments with real datasets demonstrate that compared to the baseline convert-from-blackbox-DP approach, our GP mechanisms achieve noise reductions of over an order of magnitude in several practical settings, including deep neural network training.
Yangfan Jiang 0001, Xinjian Luo, Yin Yang 0001, Xiaokui Xiao
Proc. VLDB Endow.2
2024 Exploring Privacy and Fairness Risks in Sharing Diffusion Models: An Adversarial Perspective
abstract
Diffusion models have recently gained significant attention in both academia and industry due to their impressive generative performance in terms of both sampling quality and distribution coverage. Accordingly, proposals are made for sharing pre-trained diffusion models across different organizations, as a way of improving data utilization while enhancing privacy protection by avoiding sharing private data directly. However, the potential risks associated with such an approach have not been comprehensively examined. In this paper, we take an adversarial perspective to investigate the potential privacy and fairness risks associated with the sharing of diffusion models. Specifically, we investigate the circumstances in which one party (the sharer) trains a diffusion model using private data and provides another party (the receiver) black-box access to the pre-trained model for downstream tasks. We demonstrate that the sharer can execute fairness poisoning attacks to undermine the receiver’s downstream models by manipulating the training data distribution of the diffusion model. Meanwhile, the receiver can perform property inference attacks to reveal the distribution of sensitive features in the sharer’s dataset. Our experiments conducted on real-world datasets demonstrate remarkable attack performance on different types of diffusion models, which highlights the critical importance of robust data auditing and privacy protection protocols in pertinent applications.
Xinjian Luo, Yangfan Jiang 0001, Fei Wei, Yuncheng Wu, Xiaokui Xiao, Beng Chin Ooi
IEEE Trans. Inf. Forensics Secur.1
2024 On Data Distribution Leakage in Cross-Silo Federated Learning
abstract
Federated learning (FL) has emerged as a promising privacy-preserving machine learning paradigm, enabling data owners to collaboratively train a joint model by sharing model parameters instead of private training data. However, recent studies reveal the privacy risks in FL by inferring private training data from model parameters. Therefore, differential privacy (DP) is incorporated into FL to safeguard training data. Nevertheless, DP does not provide a strong theoretical guarantee for protecting data distribution, which is also highly sensitive in thecross-siloFL scenarios as it may reflect the business secrets of data owners. In this paper, we develop two attack methods to investigate the potential risks of data distribution leakage in differentially private cross-silo FL. We highlight that an honest-but-curious server can successfully infer both the feature and label distributions of each party's training data without any background knowledge. Specifically, the first attack applies when models are differentiable, while the second attack caters to non-differentiable classification models. Extensive experiments on six benchmark datasets validate the effectiveness of the proposed attacks. The results demonstrate that the state-of-the-art DP-SGD algorithm is still vulnerable to the inference attack on data distribution, emphasizing the necessity of designing more advanced privacy-preserving FL frameworks.
Yangfan Jiang 0001, Xinjian Luo, Yuncheng Wu, Xiaochen Zhu 0003, Xiaokui Xiao, Beng Chin Ooi
IEEE Trans. Knowl. Data Eng.2
2022 A Fusion-Denoising Attack on InstaHide with Data Augmentation
abstract
InstaHide is a state-of-the-art mechanism for protecting private training images, by mixing multiple private images and modifying them such that their visual features are indistinguishable to the naked eye. In recent work, however, Carlini et al. show that it is possible to reconstruct private images from the encrypted dataset generated by InstaHide. Nevertheless, we demonstrate that Carlini et al.’s attack can be easily defeated by incorporating data augmentation into InstaHide. This leads to a natural question: is InstaHide with data augmentation secure? In this paper, we provide a negative answer to this question, by devising an attack for recovering private images from the outputs of InstaHide even when data augmentation is present. The basic idea is to use a comparative network to identify encrypted images that are likely to correspond to the same private image, and then employ a fusion-denoising network for restoring the private image from the encrypted ones, taking into account the effects of data augmentation. Extensive experiments demonstrate the effectiveness of the proposed attack in comparison to Carlini et al.’s attack.
Xinjian Luo, Xiaokui Xiao, Yuncheng Wu, Beng Chin Ooi
AAAI1
2022 Feature Inference Attack on Shapley Values
abstract
As a solution concept in cooperative game theory, Shapley value is highly recognized in model interpretability studies and widely adopted by the leading Machine Learning as a Service (MLaaS) providers, such as Google, Microsoft, and IBM. However, as the Shapley value-based model interpretability methods have been thoroughly studied, few researchers consider the privacy risks incurred by Shapley values, despite that interpretability and privacy are two foundations of machine learning (ML) models.
Xinjian Luo, Yangfan Jiang 0001, Xiaokui Xiao
CCS1
2021 Feature Inference Attack on Model Predictions in Vertical Federated Learning
abstract
Federated learning (FL) is an emerging paradigm for facilitating multiple organizations' data collaboration without revealing their private data to each other. Recently, vertical FL, where the participating organizations hold the same set of samples but with disjoint features and only one organization owns the labels, has received increased attention. This paper presents several feature inference attack methods to investigate the potential privacy leakages in the model prediction stage of vertical FL. The attack methods consider the most stringent setting that the adversary controls only the trained vertical FL model and the model predictions, relying on no background information of the attack target's data distribution. We first propose two specific attacks on the logistic regression (LR) and decision tree (DT) models, according to individual prediction output. We further design a general attack method based on multiple prediction outputs accumulated by the adversary to handle complex models, such as neural networks (NN) and random forest (RF) models. Experimental evaluations demonstrate the effectiveness of the proposed attacks and highlight the need for designing private mechanisms to protect the prediction outputs in vertical FL.
Xinjian Luo, Yuncheng Wu, Xiaokui Xiao, Beng Chin Ooi
ICDE1
2019 Accelerate Data Retrieval by Multi-Dimensional Indexing in Switch-Centric Data Centers
abstract
Data centers, receiving increased attention in data management and analysis communities, have posed new challenges in data-intensive applications, among which efficient querying processing holds a critical position. To accelerate the efficiency of multi-dimensional data retrieval, we propose a distributed multi-dimensional indexing scheme for switch-centric data centers in this paper. We first propose FR-Index, a two-layer indexing system integrating both Fat-tree topology and R-tree indexing structure. In the lower layer, each server indexes the local data with R-tree, while in the upper layer the distributed global index depicting an overview of the whole dataset. Based on the Fat-tree topology, we design a specific indexing space partitioning and mapping strategy for efficient global index maintenance and query processing. Furthermore, we develop a cost model to dynamically update FR-Index. Experiments on Amazon’s EC2 platform, comparing FR-Index with RT-CAN and RB-Index, show that the proposed indexing schema is scalable, efficient and lightweight, which can significantly promote the efficiency of query processing in data centers.
Xinjian Luo, Xiaofeng Gao 0001, Guihai Chen
Comput. J.1
2018 D2-Tree: A Distributed Double-Layer Namespace Tree Partition Scheme for Metadata Management in Large-Scale Storage Systems
abstract
The behavior of metadata server (MDS) cluster is critically important to the overall performance of today's petabyte-scale or even exabyte-scale distributed file system. How to maintain a high level of both system locality and load balancing is a significant challenge to MDS clusters. However, traditional metadata management schemes, including hash-based mapping and subtree partitioning, have severe bias on either system locality or load balancing. In this paper, we propose D2-Tree, a distributed double-layer namespace tree partition scheme, for metadata management in large-scale storage systems. The innovative idea is to design a greedy strategy to split the namespace tree into global layer and local layer subtrees, of which global layer is replicated to maintain load balancing and the lower-half subtrees are allocated separately to MDS's by a mirror division method to preserve locality. Both theoretical analysis based on empirical cumulative distribution and extensive experiments are provided to validate the efficiency of D2-Tree. Experiments using actual trace data on Amazon EC2 also exhibit the superior performance of D2-Tree compared with much previous literature.
Xinjian Luo, Xiaofeng Gao 0001, Zhaowei Tan, Xiaochun Yang 0001, Guihai Chen
ICDCS1
2018 Pipelining collaborative test for improving student performance in introductory programming courses
abstract
This study presents an innovative pipelining two-stage exam format, PipE2, to improve the first year students’ performance. Experiment results show that the impact on student performance during PipE2 was influenced by the difficulty of questions. More exposure to the topics before group discussion could improve the benefits and efficiency of PipE2.
Xinjian Luo, Qianni Deng
ITiCSE1
2010 The approximation characteristic of diagonal matrix in probabilistic setting
Guanggui Chen, Pengjuan Nie, Xinjian Luo
J. Complex.3