EDBT 2026 Demo / reviewers in the wild / expert
Xue Yang 0003
dblp:13/1779-3
· DBLP profile ↗
21ranked-venue papers
7as first author
15since 2021 · last 2026
0000-0002-4083-729XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 5 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 2 since 2021Computer networks · 4 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Practical privacy-preserving federated learning based on multiparty homomorphic encryption for large-scale models
Xian Qin, Xue Yang 0003, Xiaohu Tang 0004 |
Pattern Recognit. | 2 |
| 2026 | Constant-Round Privacy-Preserving KNN Classification Based on Function Secret SharingabstractPrivacy-preserving $k$-nearest neighbors (KNN) classification has attracted significant attention in recent years. However, existing schemes often face challenges such as high computational cost and excessive communication rounds, which limit their practical applicability. In this paper, we propose a constant-round privacy-preserving KNN classification scheme based on function secret sharing (FSS) with two non-colluding servers. To enhance data privacy and computation efficiency in secure KNN classification, we design several lightweight secure two-party computation (2PC) protocols, including Euclidean distance computation, integer comparisons, and frequency computation. To further reduce communication rounds, we introduce a batch comparison algorithm that efficiently sorts a set to extract the $k$-minimum values and the maximum value. Compared to the best-known schemes that require $\mathcal{O}(n + k \log n)$ or $\mathcal{O}(kn)$ communication rounds, where $n$ represents the dataset size, our approach achieves only 10 communication rounds. Security analysis confirms that the proposed scheme effectively preserves data privacy. Performance evaluations demonstrate that our scheme is competitive with existing works in terms of accuracy, computation cost, and communication efficiency. Bin Liu 0070, Xue Yang 0003, Xiaohu Tang 0004 |
IEEE Trans. Big Data | 2 |
| 2026 | Efficient Byzantine-Robust Privacy-Preserving Federated Learning via Dimension CompressionabstractFederated Learning (FL) allows collaborative model training across distributed clients without sharing raw data, thus preserving privacy. However, the system remains vulnerable to privacy leakage from gradient updates and Byzantine attacks from malicious clients. Existing solutions face a critical trade-off among privacy preservation, Byzantine robustness, and computational efficiency. We propose a novel scheme that effectively balances these competing objectives by integrating homomorphic encryption with dimension compression based on the Johnson-Lindenstrauss transformation. Our approach employs a dual-server architecture that enables secure Byzantine defense in the ciphertext domain while dramatically reducing computational overhead through gradient compression. The dimension compression technique preserves the geometric relationships necessary for Byzantine defence while reducing computation complexity fromO(dn)toO(kn)cryptographic operations, wherekd. Extensive experiments across diverse datasets demonstrate that our approach maintains model accuracy comparable to non-private FL while effectively defending against Byzantine clients comprising up to 40% of the network. Our approach also demonstrates substantial improvements in computational and communication efficiency. Experimental evaluation shows that the dimension compression technique achieves 25× ~ 35× reduction in computational overhead and 17× reduction in communication overhead compared to our non-compression version. When compared to state-of-the-art methods like ShieldFL [1], our approach demonstrates order-of-magnitude improvements in both computational and communication efficiency while maintaining equivalent privacy guarantees and achieving superior Byzantine robustness comparable to FLTrust [2]. These substantial efficiency enhancements make secure FL practical for deployment in large-scale neural networks with millions of parameters. Xian Qin, Xue Yang 0003, Xiaohu Tang 0004 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2024 | An efficient privacy-preserving and verifiable scheme for federated learning
Xue Yang 0003, Minjie Ma, Xiaohu Tang 0004 |
Future Gener. Comput. Syst. | 1 |
| 2024 | Efficiently Achieving Privacy Preservation and Poisoning Attack Resistance in Federated LearningabstractFederated learning enables clients to train models locally and provide local updates to the server instead of raw dataset, thereby preserving data privacy to some extent. However, adversaries can still pry users’ privacy by inferring updates, and compromise the integrity of the global model through poisoning attack. Therefore, many related works have integrated poisoning attack detection method with secure computation to address both issues. Nevertheless, they still encounter two major challenges: (i) the efficiency is too low to be applied in practice, and (ii) the privacy is still at risk of being leaked, e.g., the distance of two local updates for detecting poisoning attack could be exposed to the server. Aiming at the challenges, in this paper, we propose an Efficient Privacy-preserving and Poisoning attack Resistant scheme for Federated Learning, named EPPRFL, which preserves the privacy for local updates and some intermediate information used to detect poisoning attack. In particular, we design an efficient poisoning attack detection method based on Euclidean distance filtering & clipping technique, named F&C. Then, considering the privacy preservation of the F&C method, we efficiently customize secure comparison, secure median, secure distance computation and secure clipping protocols based on additive secret sharing. Experimental results and theoretical analysis show that compared with existing schemes, EPPRFL can better resist poisoning attack and has lower computational and communication overheads on the client side. Xue-Yang Li, Xue Yang 0003, Zhengchun Zhou, Rongxing Lu |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2024 | An Efficient and Multi-Private Key Secure Aggregation Scheme for Federated LearningabstractIn light of the emergence of privacy breaches in federated learning, secure aggregation protocols, which mainly adopt either homomorphic encryption or threshold secret sharing techniques, have been extensively developed to preserve the privacy of each client's local gradient. Nevertheless, many existing schemes suffer from either poor capability of privacy protection or expensive computational and communication overheads. Accordingly, in this paper, we propose an efficient and multi-private key secure aggregation scheme for federated learning. Specifically, we skillfully design a multi-private key secure aggregation algorithm that achieves homomorphic addition operation, with two important benefits: 1) both the server and each client can freely select public and private keys without introducing a trusted third party, and 2) the plaintext space is relatively large, making it more suitable for deep models. Besides, for dealing with the high dimensional deep model parameter, we introduce a super-increasing sequence to compress multi-dimensional data into one dimension, which greatly reduces encryption and decryption times as well as communication for ciphertext transmission. Detailed security analyses show that our proposed scheme can achieve semantic security of both individual local gradients and the aggregated result while achieving optimal robustness in tolerating client collusion. Extensive simulations demonstrate that the accuracy of our scheme is almost the same as the non-private approach, while the efficiency of our scheme is much better than the state-of-the-art baselines. More importantly, the efficiency advantages of our scheme will become increasingly prominent as the number of model parameters increases. Xue Yang 0003, Zifeng Liu, Xiaohu Tang 0004, Rongxing Lu |
IEEE Trans. Serv. Comput. | 1 |
| 2023 | Black-Box Dataset Ownership Verification via Backdoor WatermarkingabstractDeep learning, especially deep neural networks (DNNs), has been widely and successfully adopted in many critical applications for its high effectiveness and efficiency. The rapid development of DNNs has benefited from the existence of some high-quality datasets (e.g., ImageNet), which allow researchers and developers to easily verify the performance of their methods. Currently, almost all existing released datasets require that they can only be adopted for academic or educational purposes rather than commercial purposes without permission. However, there is still no good way to ensure that. In this paper, we formulate the protection of released datasets as verifying whether they are adopted for training a (suspicious) third-party model, where defenders can only query the model while having no information about its parameters and training details. Based on this formulation, we propose to embed external patterns via backdoor watermarking for the ownership verification to protect them. Our method contains two main parts, including dataset watermarking and dataset verification. Specifically, we exploit poison-only backdoor attacks (e.g., BadNets) for dataset watermarking and design a hypothesis-test-guided method for dataset verification. We also provide some theoretical analyses of our methods. Experiments on multiple benchmark datasets of different tasks are conducted, which verify the effectiveness of our method. The code for reproducing main experiments is available at https://github.com/THUYimingLi/DVBW. Yiming Li 0004, Mingyan Zhu 0001, Xue Yang 0003, Yong Jiang 0001, Tao Wei 0002, Shutao Xia |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2022 | An Accuracy-Lossless Perturbation Method for Defending Privacy Attacks in Federated LearningabstractAlthough federated learning improves privacy of training data by exchanging local gradients or parameters rather than raw data, the adversary still can leverage local gradients and parameters to obtain local training data by launching reconstruction and membership inference attacks. To defend against such privacy attacks, many noises perturbed methods (like differential privacy or CountSketch matrix) have been widely designed. However, the strong defence ability and high learning accuracy of these schemes cannot be ensured at the same time, which will impede the wide application of FL in practice (especially for medical or financial institutions that require both high accuracy and strong privacy guarantee). To overcome this issue, we propose an efficient model perturbation method for federated learning to defend against reconstruction and membership inference attacks launched by curious clients. On the one hand, similar to the differential privacy, our method also selects random numbers as perturbed noises added to the global model parameters, and thus it is very efficient and easy to be integrated in practice. Meanwhile, the random selected noises are positive real numbers and the corresponding value can be arbitrarily large, and thus the strong defence ability can be ensured. On the other hand, unlike differential privacy or other perturbation methods that cannot eliminate added noises, our method allows the server to recover the true aggregated gradients by eliminating the added noises. Therefore, our method does not hinder learning accuracy at all. Extensive experiments demonstrate that for both regression and classification tasks, our method achieves the same accuracy as non-private approaches and outperforms the state-of-the-art defence schemes. Besides, the defence ability of our method against reconstruction and membership inference attack is significantly better than the state-of-the-art related defence schemes. Xue Yang 0003, Weijun Fang, Jun Shao 0001, Xiaohu Tang 0004, Shutao Xia, Rongxing Lu |
WWW | 1 |
| 2022 | A Fine-Grained Differentially Private Federated Learning Against Leakage From GradientsabstractFederated learning (FL) enables data owners to train a global model with shared gradients while keeping private training data locally. However, recent research demonstrated that the adversary may infer private training data of clients from the exchanged local gradients, e.g., having deep leakage from gradients (DLGs). Many existing privacy-preserving approaches take usage of differential privacy (DP) to guarantee privacy. Nevertheless, the widely used privacy budget of DP (e.g., evenly distribution) leads to a sharp decline of model accuracy. To improve the model accuracy, some schemes only consider allocating the privacy budget to the fully connected layers. However, we reveal that the adversary may still reconstruct the private training data by adopting the DLG attack with the gradients of convolutional layers. In this article, we propose a fine-grained DP federated learning (DPFL) scheme, which guarantees privacy and remains high model performance simultaneously. Specifically, inspired by the methods that measure the importance of layers in deep learning, we propose a fine-grained method to allocate noise according to the importance value of layers in order to remain high model performance. Besides, we combine an active client selection strategy with DPFL and perform fine-tuning with a public data set on the server to further ensure the model performance. We evaluate DPFL under both independent and identically distributed (i.i.d) and non-i.i.d data settings to show that our method can achieve similar accuracy as the plain FL (e.g., FedAvg). We also demonstrate that our DPFL can resist the DLG attack to verify its privacy guarantee. Linghui Zhu, Yiming Li 0004, Xue Yang 0003, Shutao Xia, Rongxing Lu |
IEEE Internet Things J. | 4 |
| 2022 | Multinomial random forest
Jiawang Bai, Yiming Li 0004, Jiawei Li 0006, Xue Yang 0003, Yong Jiang 0001, Shutao Xia |
Pattern Recognit. | 4 |
| 2022 | Achieving Efficient and Privacy-Preserving Cross-Domain Big Data Deduplication in CloudabstractSecure data deduplication can significantly reduce the communication and storage overheads in cloud storage services, and has potential applications in our big data-driven society. Existing data deduplication schemes are generally designed to either resist brute-force attacks or ensure the efficiency and data availability, but not both conditions. We are also not aware of any existing scheme that achieves accountability, in the sense of reducing duplicate information disclosure (e.g., to determine whether plaintexts of two encrypted messages are identical). In this paper, we investigate a three-tier cross-domain architecture, and propose an efficient and privacy-preserving big data deduplication in cloud storage (hereafter referred to as EPCDD). EPCDD achieves both privacy-preserving and data availability, and resists brute-force attacks. In addition, we take accountability into consideration to offer better privacy assurances than existing schemes. We then demonstrate that EPCDD outperforms existing competing schemes, in terms of computation, communication and storage overheads. In addition, the time complexity of duplicate search in EPCDD is logarithmic. Xue Yang 0003, Rongxing Lu, Kim-Kwang Raymond Choo, Fan Yin, Xiaohu Tang 0004 |
IEEE Trans. Big Data | 1 |
| 2022 | Achieving Efficient Secure Deduplication With User-Defined Access Control in CloudabstractCloud storage as one of the most important services of cloud computing which significantly facilitates cloud users to outsource their data to the cloud for storage and share them with authorized users. In cloud storage, secure deduplication has been widely investigated as it can eliminate the redundancy over the encrypted data to reduce storage space and communication overhead. Regarding the security and privacy, many existing secure deduplication schemes generally focus on achieving the following properties: data confidentiality, tag consistency, access control, and resistance to brute-force attacks. However, as far as we know, none of them can achieve these four requirements at the same time. To overcome this shortcoming, in this article, we propose an efficient secure deduplication scheme that supports user-defined access control. Specifically, by allowing only the cloud service provider to authorize data access on behalf of data owners, our scheme can maximally eliminate duplicates without violating the security and privacy of cloud users. Detailed security analysis shows that our authorized secure deduplication scheme achieves data confidentiality and tag consistency while resisting brute-force attacks. Furthermore, extensive simulations demonstrate that our scheme outperforms the existing competing schemes, in terms of computational, communication and storage overheads as well as the effectiveness of deduplication. Xue Yang 0003, Rongxing Lu, Jun Shao 0001, Xiaohu Tang 0004, Ali A. Ghorbani 0001 |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2021 | GDST: Global Distillation Self-Training for Semi-Supervised Federated LearningabstractFederated Learning (FL) refers to the machine learning scheme that enables decentralized model training over massive separate data sources without privacy concerns. However, existing works rarely consider difficulty of obtaining sufficient data labels due to uncontrollable user behavior, especially in cross-device FL scenarios. In this paper, we consider semi-supervised federated learning (SSFL) setups and mainly focus on the disjoint scenario where local clients only have access to unlabeled data. By integrating self-training scheme for unlabeled data, we propose self-training loss as part of local training objective within federated learning framework. To further stablize and improve the learning process, we propose global distillation loss that utilize output logits of global model for per client-sample as supervision and also soften such distillation by temperature to obtain more discriminative information. Based on self-training and global distillation loss, combined with server-side training, we propose Global Distillation Self-Training (GDST) Federated Learning algorithm, which enables to distributedly learn a global model in the disjoint scenario of SSFL. Finally, we do sufficient ablation study to explore the role of each component of our GDST method, experimentally guarantee the interpretability. Linghui Zhu, Shutao Xia, Yong Jiang 0001, Xue Yang 0003 |
GLOBECOM | 5 |
| 2021 | Achieve efficient position-heap-based privacy-preserving substring-of-keyword query over cloud
Fan Yin, Rongxing Lu, Yandong Zheng, Jun Shao 0001, Xue Yang 0003, Xiaohu Tang 0004 |
Comput. Secur. | 5 |
| 2021 | Achieving Efficient and Privacy-Preserving Multi-Domain Big Data Deduplication in CloudabstractSecure data deduplication, as it can eliminate redundancies over encrypted data, has been widely developed in cloud storage to reduce storage space and communication overheads. Among them, the convergent encryption has been extensively adopted. However, it is vulnerable to brute-force attacks that can determine which plaintext in a message space corresponds to a given ciphertext. Many existing schemes have to sacrifice efficiency to resist brute-force attacks, especially for cross-domain deduplication, which is inevitably contrary to practical applications. Moreover, few existing schemes consider protecting the message equality information (i.e., whether two different ciphertexts correspond to an identical plaintext). To address the above challenges, in this paper, we propose an efficient and privacy-preserving big data deduplication scheme for a two-level multi-domain architecture. Specifically, by generating a random tag and a constant number of random ciphertexts for each data, our scheme not only ensures data confidentiality under multi-domain deduplication but also resists brute-force attacks. By allowing only the agent and cloud service provider to perform intra-deduplication and inter-deduplication, respectively, our scheme can protect the message equality information from disclosure as much as possible. Detailed security analysis shows that our scheme achieves privacy-preservation for both data content and the message equality information and data integrity while resisting brute-force attacks. Furthermore, extensive simulations demonstrate that our scheme significantly outperforms the existing competing schemes, especially the computational cost and the time complexity of the duplicate search. Xue Yang 0003, Rongxing Lu, Jun Shao 0001, Xiaohu Tang 0004, Ali A. Ghorbani 0001 |
IEEE Trans. Serv. Comput. | 1 |
| 2020 | SDCN: Sparsity and Diversity Driven Correlation Networks for Traffic Demand ForecastingabstractTraffic demand forecasting is essential to intelligent transportation systems and is widely used to support urban planning, traffic management and vehicle dispatching. One challenge of this problem is to model the complex spatial-temporal correlation. Although both factors have been studied, many of the existing works have strong limitations. They rely too heavily on the locality assumption (i.e., the local area is more relevant than the remote area) and only use a distance-based correlation measurement. However, the spatial correlation is also global (i.e., areas far away may also be relevant) and sparse. And it's insufficient to measure the spatial correlation using only the distance measurement. In this paper, a sparsity and diversity driven correlation network is proposed to tackle these issues. Firstly, Multiple sparse correlation graphs are carefully generated to encode sparsity and diversity. Then a newly designed hybrid graph filtering module (HGFM) leverages them to learn a more expressive node representation. Finally, the HGFM-based recurrent filtering module (RFM) is introduced to handle the spatial-temporal correlation. Extensive experiments conducted on real-world datasets demonstrate the competitiveness of our model while showing the significance of sparsity and diversity. Wenjie Li 0008, Xue Yang 0003, Xiaohu Tang 0004, Shutao Xia |
IJCNN | 2 |
| 2019 | Achieving Efficient and Privacy-Preserving Top-k Query Over Vertically Distributed Data SourcesabstractData collected from various data sources are destined to be logically interrelated but geographically distributed. Top-k query is an efficient way to find the most important objects from high volumes of data. A common way to process the top-k query over distributed data is to bring them to a centralized entity (e.g. cloud). However, there are privacy considerations during the top-k query when dealing with sensitive data (e.g. eHealthcare data) in such method. Apart from data privacy, efficiency also needs to be taken into consideration. Existing focuses on top-k query do not (fully) consider the data privacy or efficiency. In order to deal with the mentioned disadvantages, in this paper, we propose an efficient and privacy-preserving top-k query scheme over vertically distributed data. Specifically, we first design a data filtering technique to reduce the number of transmitted data from each data source to the centralized entity, which can greatly reduce the communication overhead and computational cost. Then, we propose a privacy-preserving top-k query scheme over encrypted data by deploying the homomorphic encryption technique, which can well preserve the private information and achieve the functionality at the same time. Besides, security analysis shows that the proposed scheme is privacy-preserving and performance evaluation validates the efficiency of the proposed scheme. Yandong Zheng, Rongxing Lu, Xue Yang 0003, Jun Shao 0001 |
ICC | 3 |
| 2019 | An Efficient and Privacy-Preserving Disease Risk Prediction Scheme for E-HealthcareabstractBig data mining-driven disease risk prediction has become one of the important topics in the field of e-healthcare. However, without the security and privacy assurances, disease risk prediction cannot continue to flourish. To address this challenge, in this paper, an efficient and privacy-preserving disease risk prediction scheme for e-healthcare is proposed, hereafter referred to as EPDP. Compared with the up-to-date works, the proposed EPDP comprehensively achieves two phases of disease risk prediction, i.e., disease model training and disease prediction, while ensuring the privacy preservation. Specifically, a super-increasing sequence is combined with a homomorphic cryptographic algorithm to efficiently extract the symptom set of each disease in the phase of disease model training. Bloom filter technique is introduced to compute the prediction result in the phase of disease risk prediction. Besides, extensive performance evaluations demonstrate that our proposed EPDP attains outstanding efficiency advantage over the state-of-the-art in terms of both computational and communication overheads, and hence our EPDP is more suitable for real-time e-healthcare, especially medical emergency. Xue Yang 0003, Rongxing Lu, Jun Shao 0001, Xiaohu Tang 0004, Haomiao Yang |
IEEE Internet Things J. | 1 |
| 2018 | An Improved User Identification Method Across Social Networks Via Tagging BehaviorsabstractUser Identification problem is concerned with identifying the same person with multiple virtual identities across social network sites(SNSs). Most of the existing approaches pays close attention to the similarity of profile attributes, generate-contents and linkages of friends or simply combination of these features. Only one method analyzes the feasibility of user tags in User Identification problems, but does not analyze the particularity and the inconsistency of tags belong to users among different social networks. In this paper, an improved user identification method across social networks via tagging behaviors is proposed that a new symmetric variant of BM25 (BM25 is a bag-of-words retrieval function that ranks a set of documents based on the query terms appearing in each document, regardless of the inter-relationship between the query terms within a document) using the semantic relationships between inconsistent tags among different social networks. By using extracted features from the inconsistent tagging behaviors, profile attributes and SVM supervised learning techniques, a classifier is developed for performing user identity matching between two social network sites. Evaluation on Douban and Weibo real world data-set showed that the accuracy of the proposed method is 30% higher than that of the common tag-based approach. Ning Zheng 0001, Ming Xu 0001, Xue Yang 0003, Jian Xu 0001 |
ICTAI | 4 |
| 2017 | A Privacy Settings Prediction Model for Textual Posts on Social Networks
Ming Xu 0001, Xue Yang 0003, Ning Zheng 0001, Yiming Wu 0001, Jian Xu 0001 |
CollaborateCom | 3 |
| 2017 | An Efficient Black-Box Vulnerability Scanning Method for Web Application
Haoxia Jin, Ming Xu 0001, Xue Yang 0003, Ting Wu 0001, Ning Zheng 0001 |
CollaborateCom | 3 |