VLDB 2026 Research / reviewers in the wild / expert
Xuebin Ma
dblp:27/4793
· DBLP profile ↗
24ranked-venue papers
10as first author
16since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 7 · 3 first-author · 6 since 2021Security and privacy · 7 · 3 first-author · 3 since 2021Software engineering, systems software and programming languages · 3 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | PDP-FD: Federated Knowledge Distillation Based on Personalized Differential PrivacyabstractFederated learning (FL) is a privacy-preserving distributed machine learning approach that enables training by exchanging model parameters without uploading local private data. However, the data heterogeneity across clients poses significant challenges in achieving personalized privacy protection. Existing methods often struggle to balance privacy and model utility, especially with inconsistent data distributions. Personalized differential privacy (PDP) is a commonly used technique to provide differential privacy (DP) by introducing varying levels of noise for each client. The noise level directly affects the model's utility, making it crucial to precisely determine the appropriate noise for each client. To address this challenge, we propose a federated knowledge distillation method based on PDP, named PDP-FD. PDP-FD dynamically adjusts the network architecture of both base and personalized layers to align with the data characteristics of different clients, enhancing model personalization. Moreover, it allocates an appropriate privacy budget to each client based on the similarity between local and global models, thereby meeting diverse privacy needs. Experimental results show that PDP-FD significantly outperforms existing FL methods in accuracy, effectively balancing privacy protection and model utility. Yushuang Xiao, Juanjuan Wang, Xuebin Ma, Xinwen Zhang, Xiangyu Bai |
CSCWD | 4 |
| 2025 | DPTB-VFL: An Efficient Vertical Federated Learning Framework Based on Boosting Trees and Adaptive Differential PrivacyabstractVertical Federated Learning (VFL) offers a promising approach to collaborative model training, allowing participants to share identical data samples with distinct attributes. This approach avoids direct raw data sharing, enhancing data privacy, though it may sometimes reduce model accuracy. In this paper, we introduce DPTB-VFL, a differential privacy-based vertical federated learning framework that leverages boosting trees. Boosting trees are chosen for their exceptional ability to handle heterogeneous features, deliver robust performance with minimal preprocessing, and provide interpretability-all crucial in privacy-sensitive scenarios. Additionally, most current methods apply a uniform privacy budget across all training steps, overlooking variations in local gradients that could better balance privacy and utility. DPTB-VFL addresses this gap by introducing an adaptive differential privacy protocol, enabling dynamic privacy budget allocation based on the model's learning progress. Experimental evaluations on three public datasets demonstrate that DPTB-VFL not only achieves higher accuracy than existing approaches but also significantly reduces computational latency, aligning data privacy needs with model performance Xinwen Zhang, Xuebin Ma, Yushuang Xiao, Xiangyu Bai |
CSCWD | 2 |
| 2025 | An Efficient Federated Learning with Correlation-Based Pruning: Improving Accuracy under Layer-Wise Differential PrivacyabstractFederated Learning (FL) enables multiple clients to collaboratively train models without sharing data. However, it commonly faces dual challenges of security and high communication costs. Differential Privacy (DP) offers protection by adding noise to model parameters based on strict privacy standards, but excessive noise can compromise model accuracy. Additionally, the communication cost associated with training large-scale models in FL can be both slow and expensive. In this paper, we propose CPDP-FL, an efficient and privacy-preserving federated learning algorithm that combines model pruning with differential privacy to address these issues. By pruning the model based on neuron correlation before client training, we reduce redundant parameters, which not only improves communication efficiency but also reduces the amount of noise needed for DP, thereby preserving model accuracy. During training, we apply differential privacy to the remaining parameters and introduce a novel layer-wise privacy budget allocation strategy. This approach assigns different privacy budgets to different layers to balance privacy protection with model accuracy. Extensive experiments demonstrate that our method achieves high communication efficiency and robust privacy protection while minimizing unnecessary privacy budget expenditure. Xuebin Ma, Xinwen Zhang, Yushuang Xiao, Xiangyu Bai |
CSCWD | 2 |
| 2024 | AdaDP-CFL: Cluster Federated Learning with Adaptive Clipping Threshold Differential PrivacyabstractFederated learning is a distributed machine learning approach that enables multiple clients to train models collaboratively. As its data remains stored locally on each client, this approach significantly enhances the protection of private information. However, federated learning still faces privacy leakage risks in environments with data heterogeneity. Differential privacy mechanisms are widely utilized in federated learning to ensure privacy for clients, and the magnitude of the clipping threshold directly impacts model utility. The current research does not adequately address the impact of model accuracy and training loss on clipping thresholds and is challenged by excessive hyperparameter adjustments. In response to these challenges, we propose an adaptive clipping-based differential privacy federated learning algorithm named AdaDP-CFL. It achieves model personalization and facilitates knowledge sharing among different groups through clustering and regularization techniques. Subsequently, the algorithm addresses the issue of adaptive clipping for various clients, formulated as a Markov decision process, by utilizing a deep deterministic policy gradient model based on gradient differences across client groups. Experimental results demonstrate that our algorithm outperforms current algorithms in accuracy, effectively balancing privacy protection and model utility. Xuebin Ma |
IWQoS | 2 |
| 2024 | ST-TrajGAN: A synthetic trajectory generation algorithm for privacy preservation
Xuebin Ma, Zinan Ding |
Future Gener. Comput. Syst. | 1 |
| 2024 | A privacy-preserving group decision making expert system for medical diagnosis based on dynamic knowledge base
Wuyungerile Li, Na Zong, Xuebin Ma |
Wirel. Networks | 5 |
| 2023 | Improved Bayesian network differential privacy data-releasing method based on junction tree
Xuebin Ma, Xuejian Qi, Yulei Meng |
COMPSAC | 1 |
| 2023 | RL-KDA: A K-degree Anonymity Algorithm Based on Reinforcement LearningabstractK-degree anonymity is one of the main techniques for data privacy and has gained attention in academia, industry, and government. Many social network data publishing algorithms based on K-anonymity techniques have been proposed, but most studies focus on static social networks. Compared to static social networks, dynamic social networks suffer from problems such as higher information loss and lower data utility. To address the existing problem of dynamic social networks, we propose a K-degree anonymity dynamic data publishing algorithm based on reinforcement learning. The algorithm ends with two phases: anonymization sequence and graph modification. In the anonymous sequence phase, this paper combines the idea of reinforcement learning and the characteristics of dynamic data change to build a reinforcement learning model for anonymous sequences. In this way, an ideal anonymous sequence can be created. We also propose a new strategy for graph modification, which selects edges according to degree centrality to generate anonymous graphs. Finally, experiments on real datasets show the effectiveness of our algorithm. Xuebin Ma, Yulan Gao |
COMPSAC | 1 |
| 2023 | Differential Privacy Frequent Closed Itemset Mining over Data StreamabstractAs the most common technique for mining and analyzing massive data, frequent itemset mining is widely used in various scenarios. However, when the data contains sensitive information, it will bring serious privacy leakage risks to mining and publishing it directly. Therefore, how to efficiently mine frequent itemset without privacy disclosure is a hot issue at present. In data stream, because the data between adjacent timestamps possess certain correlation, which makes it easier to leak privacy for frequent itemset mining in the data stream, and considering that frequent itemset will have combinatorial explosion problem in data stream, frequent closed itemset mining in the data stream will be a better choice. In this paper, we propose a differential privacy algorithm for mining frequent closed itemset in data stream, which is referred to as DPES. The algorithm uses a vertical mining method in each sliding window, which can quickly obtain the frequent closed itemset contained in the current window. Besides, to improve the data utility, we also design an adaptive privacy budget allocation strategy by computing the difference to decide whether the current timestamp should publish a low-noise statistic or an approximate statistic. Finally, we demonstrate that the DPES algorithm satisfies ε-difference privacy by privacy analysis, and the experimental results on several real datasets also show the effectiveness of the DPES algorithm. Xuebin Ma, Shengyi Guan, Yanan Lang |
TrustCom | 1 |
| 2023 | Unsupervised Domain Adaptation with Differentially Private Gradient ProjectionabstractDomain adaptation is a viable solution for deep learning with small data. However, domain adaptation models trained on data with sensitive information may be a violation of personal privacy. In this article, we proposed a solution for unsupervised domain adaptation, called DP‐CUDA, which is based on differentially private gradient projection and contradistinguisher. Compared with the traditional domain adaptation process, DP‐CUDA involves searching for domain‐invariant features between the source domain and target domain first and then transferring knowledge. Specifically, the model is trained in the source domain by supervised learning from labeled data. During the training of the target model, feature learning is used to solve the classification task in an end‐to‐end manner using unlabeled data directly, and the differentially private noise is injected into the gradient. We conducted extensive experiments on a variety of benchmark datasets, including MNIST, USPS, SVHN, VisDA‐2017, Office‐31, and Amazon Review, to demonstrate our proposed method’s utility and privacy‐preserving properties. Maobo Zheng, Xuebin Ma |
Int. J. Intell. Syst. | 3 |
| 2022 | TKDA: An Improved Method for K-degree Anonymity in Social GraphsabstractData anonymization is one of the most important directions in privacy-preserving. However, research shows that simple anonymization of data does not protect privacy. To solve this problem, we present a novel and effective algorithm named tree-based K-degree anonymity (TKDA). We devise a new anonymity sequence generation method to reduce the information loss for social graphs. Then, the dynamic anonymization process is implemented by a depth-first search (DFS) traversal algorithm. Finally, the graph modification algorithm based on the anonymous sequence can keep the original graph structure stable. Average Path Length (APL), Average Clustering Coefficient (ACC), and Transitivity (T) are employed to evaluate the method. Experimental results on several datasets show that TKDA is closer to the values of the original graphs on the correlated three experimental metrics, which indicates that TKDA portrays the real data in more detail and improves the utility of the released data. Xuebin Ma |
ISCC | 2 |
| 2022 | Publishing Weighted Graph with Node Differential PrivacyabstractAt present, how to protect user privacy and security while publishing user data has become an increasingly important problem. Differential privacy is mainly divided into two directions in graph data publishing. One is to publish the statistical characteristics of the graph that meets the differential privacy, and the other is to publish the synthesis graph that meets the differential privacy. This paper proposes a weighted graph publishing method based on node difference privacy. First, this paper proposes a projection method that constrains the degree of nodes and the number of triangles and reduces the increase in noise by reducing the sensitivity. Afterward, select appropriate statistical characteristics of the weighted graph to form node attributes as the parameters of the syn-thesis weighted graph. The next part proposes a graph publishing method based on node attributes and weights. This method synthesizes the initial graph according to the degree in the node attribute. It then adds or deletes the edges of the initial graph according to the number of triangles in the node attribute to obtain the final synthesis graph. Finally, this paper verifies the weighted graph publishing method proposed on three data sets. The results show that the method proposed in this paper satisfies the different privacy conditions of nodes while maintaining certain utility. Xuebin Ma, Ganghong Liu, Aixin Lin |
MSN | 1 |
| 2022 | RDP-WGAN: Image Data Privacy Protection Based on Rényi Differential PrivacyabstractIn recent years, artificial intelligence technology based on image data has been widely used in various industries. Rational analysis and mining of image data can not only promote the development of the technology field but also become a new engine to drive economic development. However, the privacy leakage problem has become more and more serious. To solve the privacy leakage problem of image data, this paper proposes the RDP-WGAN privacy protection framework, which deploys the Rényi differential privacy (RDP) protection techniques in the training process of generative adversarial networks to obtain a generative model with differential privacy. This generative model is used to generate an unlimited number of synthetic datasets to complete various data analysis tasks instead of sensitive datasets. Experimental results demonstrate that the RDP-WGAN privacy protection framework provides privacy protection for sensitive image datasets while ensuring the usefulness of the synthetic datasets. Xuebin Ma, Maobo Zheng |
MSN | 1 |
| 2022 | PU_Bpub: High-Dimensional Data Release Mechanism Based on Spectral Clustering with Local Differential Privacy
Aixin Lin, Xuebin Ma |
WASA (2) | 2 |
| 2021 | Improving the Effect of Frequent Itemset Mining with Hadamard Response under Local Differential PrivacyabstractFrequent itemset mining is a basic data mining task and has many applications in other data mining tasks. However, it is likely that the user's personal privacy information may be leaked in the mining process. In recent years, applying the local differential privacy protection model to mine all frequent itemsets is a relatively reliable and secure protection method. In local differential privacy, users first perturb the original data then send it to the aggregator, which prevents the aggregator leaking user's private information throughout the process. There are two major problems with data mining using local differential privacy, one is that the accuracy of the results is relatively low after mining, and the other is that the user transmits too much data to the server, which results in higher communication costs. In this paper, we use Hadamard response algorithm to improve the accuracy of the results while reduce the communication cost. Finally, we use FP-tree for frequent itemset mining to compare the Hadamard response with previous algorithms. Xuebin Ma, Shengyi Guan |
TrustCom | 1 |
| 2021 | DP-gSpan: A Pattern Growth-based Differentially Private Frequent Subgraph Mining AlgorithmabstractFrequent subgraph mining has been one of the researches focuses in the field of data mining, which plays an important role in understanding social interaction mechanisms, urban planning, and studying the spread of diseases in social networks. Frequent subgraph mining provides a lot of valuable information. However, mining and publishing frequent subgraphs brings more and more risks of privacy leakage. In order to solve the problem of privacy leakage, the combination of frequent subgraph mining based on Apriori and differential privacy has become the mainstream method. Nevertheless, most of the existing research studies suffer from too large candidate subgraph sets, low accuracy, and low efficiency. Therefore, we propose a more secure and effective depth-first search algorithm for frequent subgraph mining under differential privacy, which is referred to as DP-gSpan. We design a heuristic truncation strategy and a new privacy budget allocation strategy to realize the reduction of the candidate set size and the rational allocation of the privacy budget. Through privacy analysis, we prove that DP-gSpan satisfies$\varepsilon$-differential privacy. Experimental results over a large number of real-world datasets prove that the performance of the proposed mechanism is better. Jiangna Xing, Xuebin Ma |
TrustCom | 2 |
| 2020 | Differentially Private Social Graph Publishing for Community Detection
Xuebin Ma, Shengyi Guan |
SecureComm (2) | 1 |
| 2020 | DP-Eclat: A Vertical Frequent Itemset Mining Algorithm Based on Differential PrivacyabstractFrequent itemset mining has been a focused theme in the field of data mining, which is widely used in business decision making, economics, medicine, bioinformatics and other fields. Frequent itemset mining can provide a lot of valuable information when making decisions, but it may bring the risk of privacy disclosure when mining and publishing frequent itemsets. In order to solve the privacy leakage problem, most of the existing solutions are using horizontal mining method to mine frequent itemsets under differential privacy. However, these solutions generally suffer from complex support computation and poor accuracy due to large candidate sets. In this paper, we propose a new vertical frequent itemset mining algorithm based on differential privacy, which is referred to as DP-Eclat. In DP-Eclat, a new privacy budget allocation strategy is proposed to rationalize the privacy budget allocation, which allows privacy budget to be used more fully. In addition, we devise a multiple pruning strategy to further improve the data utility by prune before and after the generation of candidate itemsets. Through privacy analysis, we prove that DP-Eclat satisfies E -differential privacy. Extensive experiment results on multiple real datasets show that DP-Eclat significantly outperforms state-of-the-art algorithms in terms of data utility. Shengyi Guan, Xuebin Ma, Wuyungerile Li, Xiangyu Bai |
TrustCom | 2 |
| 2020 | Differential privacy preserving data publishing based on Bayesian networkabstractPrivacy-preserving data publishing is a hot issue in the field of privacy protection. Differential privacy is a burgeoning technology of privacy protection which provides a powerful privacy mechanism and does not make restrictive assumptions on the attacker's background knowledge. At present, there does not have an effective way to generate Synthetic high-dimensional data with differential privacy technology. Aiming at the issue of high-dimensional privacy data publishing, this paper proposed a method called APrivBayes, which altered the structure of Bayesian network to make it adapting differential privacy mechanism. Then proposed a first node selection mechanism based on attribute correlation degree for the new structure of Bayesian network. Through theoretical analysis and experimental evaluation, this method improves the effect of network and reduces Laplace noise effectively, while protecting personal privacy and improving the usability of published data. Xuejian Qi, Xuebin Ma, Xiangyu Bai, Wuyungerile Li |
TrustCom | 2 |
| 2020 | Differential Privacy Images Protection Based on Generative Adversarial NetworkabstractIn recent years, as image data are widely used in data analysis tasks, the problem of privacy disclosure is becoming more and more serious. However, the privacy protection technology of image data is still immature. In this paper, we propose a privacy protection framework named dp-WGAN for image data. This framework uses differential privacy and generative adversarial network to train a generative model with privacy protection function. Using this generative model, synthetic data with similar characteristics to sensitive data can be obtained, and synthetic data is published instead sensitive data to complete all kinds of data analysis tasks. Through extensive empirical evaluation on benchmark datasets, we demonstrate that dp-WGAN can provide strong privacy protection for sensitive data and produce high-quality synthetic data. Xuebin Ma, Xiangyu Bai, Xiangdong Su |
TrustCom | 2 |
| 2019 | Dynamic Data Publishing with Differential Privacy via Reinforcement LearningabstractDifferential privacy, which is due to its rigorous mathematical proof and strong privacy guarantee, has become a standard for the release of statistics with privacy protection. Recently, a lot of dynamic data publishing algorithms based on differential privacy have been proposed, but most of the algorithms use a native method to allocate the privacy budget. That is, the limited privacy budget is allocated to each time point uniformly, which may result in the privacy budget being unreasonably utilized and reducing the utility of data. In order to make full use of the limited privacy budget in the dynamic data publishing and improve the utility of data publishing, we propose a dynamic data publishing algorithm based on reinforcement learning in this paper. The algorithm consists of two parts: privacy budget allocation and data release. In the privacy budget allocation phase, we combine the idea of reinforcement learning and the changing characteristics of dynamic data, and establish a reinforcement learning model for the allocation of privacy budget. Finally, the algorithm finds a reasonable privacy budget allocation scheme to publish dynamic data. In the data release phase, we also propose a new dynamic data publishing strategy to publish data after the privacy budget is exhausted. Extensive experiments on real datasets demonstrate that our algorithm can allocate the privacy budget reasonably and improve the utility of dynamic data publishing. Ruichao Gao, Xuebin Ma |
COMPSAC (1) | 2 |
| 2019 | Reliable Energy-Aware Routing Protocol in Delay-Tolerant Mobile Sensor NetworksabstractIn the data transmission process of delay-tolerant mobile sensor networks, data is easily lost, and the network lifetime decreases due to energy depletion by the nodes. We propose a reliable energy-aware routing protocol, called RER. To ensure the reliability of message transmission, a hop-by-hop retransmission acknowledgement mechanism is introduced in the RER. Second, we design a metric called Reliable Energy Cost Based on Distance (RECBD) to aid RER, which is determined by analysing the distance between the current node and the relay node, the distance between the relay node and the sink node, the current residual energy of the current node, and the link quality. Finally, the message is routed based on the RECBD to improve reliability and reduce energy consumption. The simulation results show that the routing protocol can improve the energy utilization of the sensor nodes and prolong the network lifetime while guaranteeing the delivery ratio and reliability. Xuebin Ma |
Wirel. Commun. Mob. Comput. | 1 |
| 2015 | Dynamic Web Service Composition Based on State Space SearchingabstractWeb service composition problem was considered as a planning problem by previous research. However, many factors constantly affect the QoS and results of invocation of web services, thus the environment of web services is dynamic. As result, web service composition problem should be considered as an uncertain planning problem. This paper uses Markov property to deal with the uncertain planning problem for service composition. According to the uncertainty model, we propose a reinforcement learning method to compose web services. Without knowing the transition function and reward function, our uncertain planning method uses an estimated value function to approach a real function and is able to obtain a composite service. The results of experiments show that our method can effectively reduce computing time of the service composition. Yu Lei 0007, Jiantao Zhou 0002, Yongqiang Gao, Liu Jing, Xuebin Ma |
ICPADS | 5 |
| 2009 | Structural analysis of dialects, sub-dialects and sub-sub-dialects of ChineseabstractIn China, there are hundred kinds of dialects. By traditional dialectology, they are classified into seven big dialect regions and most of them also have many sub-dialects and sub-subdialects. As they are different in various linguistic aspects, people from different dialect regions often cannot communicate orally. But for the sub-dialects of one dialect region, although they are sometimes still mutually unintelligible, more common features are shared. In this paper, a dialect pronunciation structure, which has been used successfully in dialectbased speaker classification in our previous work [1], is examined for the task of speaker classification and distance measurement among cities based on sub-dialects of Mandarin. Using the finals of the dialectal utterances of a specific list of written characters, a dialect pronunciation structure is built for every speaker in a data set and these speakers are classified based on the distances among their structures. Then, the results of classifying 16 Mandarin speakers based on their sub-dialects show that they are linguistically classified with little influence of their age and gender. Finally, distances among sub-sub-dialects are similarly calculated and evaluated. All the results show high validity and accordance to linguistic studies. Index Terms: Sub-dialects of Mandarin, pronunciation structure, speaker classification Xuebin Ma, Akira Nemoto, Nobuaki Minematsu, Yu Qiao 0001, Keikichi Hirose |
INTERSPEECH | 1 |