VLDB 2026 Research / reviewers in the wild / expert
Shuo Wang 0026
dblp:63/1591-26
· DBLP profile ↗
16ranked-venue papers
6as first author
15since 2021 · last 2025
0009-0006-5388-2774ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Computer networks · 4 · 1 first-author · 4 since 2021Security and privacy · 4 · 3 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Where Does This Data Come From? Enhanced Source Inference Attacks in Federated LearningabstractFederated learning (FL) enables collaborative model training without exposing raw data, offering a privacy-aware alternative to centralized learning. However, FL remains vulnerable to various privacy attacks that exploit shared model updates, including membership inference, property inference, and gradient inversion. Source inference attacks further threaten FL by identifying which client contributed a specific training sample, posing severe risks to user and institutional privacy. Existing source inference attacks mainly assume passive adversaries and overlook more realistic scenarios where the server actively manipulates the training process. In this paper, we present an enhanced source inference attack that demonstrates how a malicious server can amplify behavioral differences between clients to more accurately infer data origin. Our approach introduces active training manipulation and data augmentation to expose client-specific patterns. Experimental results across five representative FL algorithms and multiple datasets show that our method significantly outperforms prior passive attacks. These findings reveal a deeper level of privacy vulnerability in FL and call for stronger defense mechanisms under active threat models. Xiaolong Xu 0001, Xiaokang Zhou, Fei Dai 0002, Yansong Gao 0001, Shuo Wang 0026, Hongsheng Hu |
IJCAI | 8 |
| 2025 | Fine-Grained and Efficient Self-Unlearning with Layered IterationabstractAs machine learning models become widely deployed in data-driven applications, ensuring compliance with the 'right to be forgotten' as required by many privacy regulations is vital for safeguarding user privacy. To forget the given data, existing re-labeling based unlearning methods employ a single-step adjustment scheme that revises the decision boundaries in one re-labeling phase. However, such single-step approaches lead to coarse-grained changes in decision boundaries among the remaining classes and impose adverse effects on the model utility. To address these limitations, we propose 'Self-Unlearning with Layered Iteration (SULI),' a novel unlearning approach that introduces a layered iteration strategy to re-label the forgetting data iteratively and refine the decision boundaries progressively. We further develop a 'Selective Probability Adjustment (SPA)' technique, which uses a soft-label mechanism to promote smoother decision-boundary transitions. Comprehensive experiments on three benchmark datasets demonstrate that SULI achieves superior performance in effectiveness, efficiency, and privacy compared to the state-of-the-art baselines in both class-wise and instance-wise unlearning scenarios. The source code is released at https://github.com/Hongyi-Lyu-MQ/SULI. Hongyi Lyu, Xuyun Zhang, Hongsheng Hu, Shuo Wang 0026, Chaoxiang He, Lianyong Qi |
IJCAI | 4 |
| 2025 | Differentially Private Vertical Federated Learning with Dual-Sparsification
Keke Gai, Jing Yu 0007, Shuo Wang 0026, Zhengkang Fang |
WASA (2) | 4 |
| 2025 | RAFLS: RDP-Based Adaptive Federated Learning With Shuffle ModelabstractFederated Learning (FL) realizes distributed machine learning training via sharing model updates rather than raw data, thus ensuring data privacy. However, an attacker may infer the client's local original data from the model parameter so that original data leakage can be caused. While Differential Privacy (DP) is designed to address data leakage issues in FL, injecting noises during training reduces model accuracy. To minimize the negative impact caused by noises on model accuracy while considering privacy protections, in this article we propose an adaptive FL model, entitledRDP-basedAdaptiveFederatedLearning inShuffle model (RAFLS). To ensure the dataset privacy of clients, we inject adaptive noises into the client's local model by leveraging the adaptive layer-wise adaptive sensitivity of the local model. Our approach shuffles all local model parameters in order to address privacy explosion concerns caused by high-dimensional aggregation and multiple iterations. We further propose a fine-grained model weight aggregation scheme to aggregate all local models and obtain a global model. Our experiment evaluations demonstrate the proposed RAFLS has a better performance than the state-of-the-art methods in reducing noise's impact on model accuracy while protecting data, i.e., showing that the accuracy of RAFLS increases by 1.54% than that of the baseline scheme when$\epsilon = 2.0$and FashionMNIST under IID setting. Shuo Wang 0026, Keke Gai, Jing Yu 0007, Liehuang Zhu, Hanghang Wu, Changzheng Wei, Ying Yan 0002, Hui Zhang 0002, Kim-Kwang Raymond Choo |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2025 | PraVFed: Practical Heterogeneous Vertical Federated Learning via Representation LearningabstractVertical federated learning (VFL) provides a privacy-preserving method for machine learning, enabling collaborative training across multiple institutions with vertically distributed data. Existing VFL methods assume that participants passively gain local models of the same structure and communicate with active pary during each training batch. However, due to the heterogeneity of participating institutions, VFL with heterogeneous models for efficient communication is indispensable in real-life scenarios. To address this challenge, we propose a new VFL method called Practical Heterogeneous Vertical Federated Learning via Representation Learning (PraVFed) to support the training of parties with heterogeneous local models and reduce communication costs. Specifically, PraVFed employs weighted aggregation of local embedding values from the passive party to mitigate the influence of heterogeneous local model information on the global model. Furthermore, to safeguard the passive party’s local sample features, we utilize blinding factors to protect its local embedding values. To reduce communication costs, the passive party performs multiple rounds of local pre-model training while preserving label privacy. We conducted a comprehensive theoretical analysis and extensive experimentation to demonstrate that PraVFed reduces communication overhead under heterogeneous models and outperforms other approaches. For example, when the target accuracy is set at 60% under the CINIC10 dataset, the communication cost of PraVFed is reduced by 70.57% compared to the baseline method. Our code is available athttps://github.com/wangshuo105/PraVFed_main. Shuo Wang 0026, Keke Gai, Jing Yu 0007, Zijian Zhang 0001, Liehuang Zhu |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2025 | EASTER: Embedding Aggregation-Based Heterogeneous Models Training in Vertical Federated LearningabstractVertical Federated Learning (VFL) allows collaborative machine learning without sharing local data. However, existing VFL methods face challenges when dealing with heterogeneous local models among participants, which affects optimization convergence and generalization of participants' local knowledge aggregation. To address this challenge, this paper proposes a novel approach calledEmbeddingAggregation-based HeterogeneousModelsTraining in Vertical Federated Learning(EASTER). EASTER focuses on aggregating the local embeddings of each participant's knowledge during forward propagation. We propose an embedding protection method based on lightweight blinding factors, which injects the blinding factors into the local embedding of the passive party. However, the passive party does not own the sample labels, so the local model's gradient cannot be calculated locally. To overcome this limitation, we propose a new method in which the active party assists the passive party in computing its local heterogeneous model gradients. Theoretical analysis and extensive experiments demonstrate that EASTER can simultaneously train multiple heterogeneous models and outperform some recent methods in model performance. For example, compared with the state-of-the-art method, the model accuracy of EASTER was improved by 7.22% under the CIFAR-10 dataset. Shuo Wang 0026, Keke Gai, Jing Yu 0007, Liehuang Zhu, Weizhi Meng 0001, Bin Xiao 0001 |
IEEE Trans. Mob. Comput. | 1 |
| 2024 | A Trustworthy Biological Assets Governance System Using Decentralized IdentityabstractAlong with the development of the biological industry, the system of biological assets governance has a higher-level requirement in asset identity verification due to the demands of digitization and Financial Technology (FinTech). In order to achieve broad scalability and adaptability in the governance of biological assets' identities, this work proposes a Biological Assets Decentralized Identity Management (BA-DID) system that develops a Decentralized Identity (DID)-based solution to addressing the verification issues in biological assets while considering multi-dimensional requirements. Blockchain technology is the fundamental infrastructure of the system. Authenticated certificates of the asset owners verify identity attributes. Financial Service Institutions (FSI) and governance agencies act as issuers, providing Verifiable Credentials (VC) for biological assets by using a group of attributes over a consensus. The asset owners hold VCs and can be available to other organizations as proof of the corresponding asset attributes. Our evaluations have demonstrated that the proposed approach has superior performance in biological assets verification, security, and maintenance. Zhengkang Fang, Jing Yu 0007, Shufen Fang, Shuo Wang 0026, Weilin Chan, Zexin Gao, Keke Gai |
CSCloud | 4 |
| 2024 | LRPAFL: Layer-Wise Relevance Propagation-Based Adaptive Federated LearningabstractFederated learning realizes distributed machine learning training by sharing the model rather than sharing the local dataset. However, the local dataset may be leaked during model training. While differential privacy techniques can mitigate privacy leakage to some extent, the noise tends to have a significant negative impact on model accuracy. To minimize the impact of noise on model accuracy and protect the privacy of the original data, we propose an Layer-wise Relevance Propagation-based Adaptive Federated Learning (LRPAFL). To ensure local data privacy, we inject adaptive noises that satisfy DP into the training sample according to the correlation between local training data features and the model. Specifically, we set a correlation boundary ct. We only inject an adaptive amount of noise when the correlation between the feature and the model is greater than or equal to ct. Furthermore, to evaluate the performance of our approach, we propose a relationship between privacy budget and accuracy. We theoretically and experimentally analyze the performance of this model. Compared with the baseline method, our method has a better performance and the proposed model reduces the impact of noise on model accuracy while protecting data. For example, compared with the state-of-the-art scheme, the accuracy of RASFL is increased by 2% when$\epsilon=1$ Shuo Wang 0026, Zhengkang Fang, Keke Gai |
CSCloud | 1 |
| 2024 | Shadow-Free Membership Inference Attacks: Recommender Systems Are More Vulnerable Than You Thought
Xiaoxiao Chi, Xuyun Zhang, Yan Wang 0002, Lianyong Qi, Amin Beheshti, Xiaolong Xu 0001, Kim-Kwang Raymond Choo, Shuo Wang 0026, Hongsheng Hu |
IJCAI | 8 |
| 2024 | ReVFed: Representation-Based Privacy-Preserving Vertical Federated Learning with Heterogeneous Models
Shuo Wang 0026, Jing Yu 0007, Keke Gai, Liehuang Zhu |
KSEM (3) | 1 |
| 2024 | RAC-Chain: An Asynchronous Consensus-based Cross-chain Approach to Scalable Blockchain for MetaverseabstractThe metaverse, as an emerging technical term, conceptually aims to construct a virtual digital space that runs parallel to the physical world. Due to human behaviors and interactions being represented in the virtual world, security in the metaverse is a challenging issue in which the traditional centralized service model is one of the threat sources. To conquer the obstacle caused by centralized computing, blockchain-based solutions are potential problem-solving methods. However, it is difficult for a single blockchain to support large-scale data and business services in the metaverse, due to the scalability restrictions. Moreover, multi-chain settings also encounter the interoperability issues. In this work, we propose a Relay chain and Asynchronous consensus-based Consortium blockchain cross-Chain model, which realizes message transmission and cross-chain transactions in multiple chains by adopting the relay chain and cross-chain gateways. All nodes of the application chains and the relay chain execute cross-chain transactions in sequence and reach a consensus on transactions at any transmission delay. Our experiment evaluations demonstrate that our approach performs well in atomicity, security, and functionality (cross-chain transactions), such that the performance of blockchain scalability in the metaverse can be improved, compared with the traditional relay chain schemes. Tianxiu Xie, Keke Gai, Liehuang Zhu, Shuo Wang 0026, Zijian Zhang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2023 | BDVFL: Blockchain-based Decentralized Vertical Federated LearningabstractVertical Federated Learning (VFL) effectively addresses the issue of data isolation, which makes data mining secure. Most VFL implementations rely on a single server or third party for training, which will be terminated if the server or third party fails. In addition, the model accuracy trained by VFL depends on the quality of the client’s local features; nevertheless, the client’s local feature quality is difficult to verify. There exists a chance that the features owned by the client are irrelevant to the model or the intermediate results submitted by the client are inaccurate, such that the model’s accuracy will be seriously affected. In order to solve the single point failure and model accuracy issues in VFL, this paper first proposes a Blockchain – based Decentralized VFL (BDVFL) training model. With the integration of blockchain and the VFL training process, the nodes within the blockchain are categorized into non-training and training nodes. Our method focuses on the scenario in which all training nodes possess labeled data and actively engage in the training procedure of VFL. To be specific, first, each client utilizes local features and initial models to carry out forward activation and generate intermediate results. Second, we randomly choose a training node and combine it with the intermediate results from all clients to formulate the loss function. Finally, each client updates the local model by using the gradient. To protect the raw features, a blinding factor is utilized for safeguarding the intermediate results submitted by the client, such that the training nodes cannot infer the local features from intermediate results. To mitigate the interference of irrelevant training outcomes from clients on the model’s accuracy, we propose a verifiable aggregation method to assess the validity of the intermediate results submitted by the clients. We have conducted both theoretical and experimental analysis, and the results demonstrate the effectiveness of the proposed method. Shuo Wang 0026, Keke Gai, Jing Yu 0007, Liehuang Zhu |
ICDM | 1 |
| 2023 | Blockchain-Based Multisignature Lock for UAC in MetaverseabstractAs an emerging digital concept offering interconnections across multiple platforms, the metaverse provides digital transformations for various aspects of the physical world, facilitated by a few novel technologies, for example, cloud computing offers data support for the digital world. Humans immersed in the metaverse are digital entities who communicate with others or objects, such that ubiquitous access controls (UACs) are indispensable sectors for multiple platforms. However, in the metaverse, UACs have opened a wide scope of bridges for individuals to shuttle the virtual world, which implies that numerous threats exist at the access layer due to a great pool of entries. In this paper, to solve security issues in the UAC setting of the metaverse, we propose a novel blockchain-based multisignature lock for UAC (BMSL-UAC) scheme. All data institutions reconstruct a consortium blockchain system. In addition, our proposed scheme ensures that only authorized users can access an institution’s data. Finally, we abstract the user’s data access behaviors into the transaction information of the consortium blockchain system to realize full life-cycle data management and traceability. To verify the performance of our scheme, a series of experiments are carried out on the Hyperledger, and evaluation results have demonstrated that the resource consumption, delay, and throughput of this scheme are all within a reasonable range. Keke Gai, Shuo Wang 0026, Hui Zhao 0002, Yufeng She, Zijian Zhang 0001, Liehuang Zhu |
IEEE Trans. Comput. Soc. Syst. | 2 |
| 2023 | Blockchain-Based Privacy-Preserving Positioning Data Sharing for IoT-Enabled Maritime Transportation SystemsabstractData-driven applications play an important role in modern-time maritime transportation systems, for instance in facilitating decision-making relating to communication and safety. One example application is position data sharing between vessels within the maritime Internet of Things (IoT)-enabled context. When designing such applications, we need to also consider how to ensure data accuracy as well as privacy in a large scale deployment. In this paper, we demonstrate the potential of using blockchain to facilitate privacy-preserving data sharing. Specifically, we develop a zero-knowledge proof-based scheme to protect vessel identities while allowing data sharing, and a commitment-based approach to ensure relationship-related privacy in data trading between participants. Our security and performance evaluations demonstrate the utility of the proposed approach. Keke Gai, Haokun Tang, Guangshun Li, Tianxiu Xie, Shuo Wang 0026, Liehuang Zhu, Kim-Kwang Raymond Choo |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2021 | Concurrent Practical Byzantine Fault Tolerance for Integration of Blockchain and Supply ChainabstractCurrently, the integration of the supply chain and blockchain is promising, as blockchain successfully eliminates the bullwhip effect in the supply chain. Generally, concurrent Practical Byzantine Fault Tolerance (PBFT) consensus method, named C-PBFT, is powerful to deal with the consensus inefficiencies, caused by the fast node expansion in the supply chain. However, due to the tremendous complicated transactions in the supply chain, it remains challenging to select the credible primary peers in the concurrent clusters. To address this challenge, the peers in the supply chain are classified into several clusters by analyzing the historic transactions in the ledger. Then, the primary peer for each cluster is identified by reputation assessment. Finally, the performance of C-PBFT is evaluated by conducting experiments in Fabric. Xiaolong Xu 0001, Xiaoxian Yang, Shuo Wang 0026, Lianyong Qi, Wan-Chun Dou |
ACM Trans. Internet Techn. | 4 |
| 2020 | Privacy-Enhanced Data Collection Based on Deep Learning for Internet of VehiclesabstractThe development of smart cities and deep learning technology is changing our physical world to a cyber world. As one of the main applications, the Internet of Vehicles has been developing rapidly. However, privacy leakage and delay problem for data collection remain as the key concerns behind the fast development of the cyber intelligence technologies. If the original data collected are directly uploaded to the cloud for processing, it will bring huge load pressure and delay to the network communication. Moreover, during this process, it will lead to the leakage of data privacy. To this end, in this article we design a data collection and preprocessing scheme based on deep learning, which adopts the semisupervised learning algorithm of data augmentation and label guessing. Data filtering is performed at the edge layer, and a large amount of similar data and irrelevant data are cleared. If the edge device cannot process some complex data independently, it will send the processed and reliable data to the cloud for further processing, which maximizes the protection of user privacy. Our method significantly reduces the amount of data uploaded to the cloud, and meanwhile protects the user's data privacy effectively. Tian Wang 0001, Zhihan Cao, Shuo Wang 0026, Jianhuang Wang, Lianyong Qi, Anfeng Liu, Mande Xie, Xiaolong Li 0004 |
IEEE Trans. Ind. Informatics | 3 |