VLDB 2026 Research / reviewers in the wild / expert
Wei Xu 0042
dblp:32/1213-42
· DBLP profile ↗
12ranked-venue papers
8as first author
12since 2021 · last 2026
0000-0001-8948-9516ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 8 · 6 first-author · 8 since 2021Computer networks · 3 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Kangaroo: A Private and Amortized Inference Framework over WAN for Large-Scale Decision Tree Evaluation
Wei Xu 0042, Hui Zhu 0001, Yandong Zheng, Song Bian 0001, Dengguo Feng, Hui Li 0006 |
NDSS | 1 |
| 2026 | Collusion-Resistant Privacy-Preserving Outsourced Training Under Single Cloud With Semi-Honest TEE
Wei Xu 0042, Hui Zhu 0001, Guozhang He, Jingqiang Lin 0001, Jiaqi Zhao 0005, Rongxing Lu, Dengguo Feng |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2026 | TCKKS: An Efficient TEE-Assistance CKKS Scheme Without BootstrappingabstractFully homomorphic encryption (FHE) is a powerful technique that allows unlimited computations on encrypted data without decryption. However, FHE will incur huge computation and storage costs, making it difficult to be applied in real environments. To improve the efficiency of FHE, some optimized schemes have been proposed based on the trusted execution environment (TEE), which offer a lighter and lower overhead solution for FHE optimizations to a certain extent. However, they heavily rely on the confidentiality of the TEE, and their performance is still limited. To solve the above problems, we propose an efficient TEE-assistance CKKS scheme without bootstrapping, named TCKKS, which has the characteristics of security, efficiency, low memory, and scalability. First, to weaken the trust assumption of TEE, we consider TEE to be honest-but-curious, meaning the enclave's algorithm provider will execute the algorithm honestly but might monitor the data in the enclave. Based on this assumption, we design a lightweight secure multiplication protocol (SMP) and a secure rotation protocol (SRP) for TCKKS to efficiently achieve ciphertext multiplication and rotation operations. Then, to further improve the performance of TCKKS, we optimize the arithmetic operations and encryption/decryption operations based on the characteristics of our protocols. Moreover, we prove the security of SMP and SRP under the simulation-based real/ideal worlds model and further demonstrate the security of TCKKS based on the RLWE problem. In addition, extensive experiments indicate that TCKKS has a better performance than mainstream libraries, such as RNS-HEAAN, PALISADE and SEAL. Wei Xu 0042, Hui Zhu 0001, Fengwei Wang, Yandong Zheng, Rongxing Lu, Yier Jin, Dengguo Feng |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2026 | Breaking Beyond One: Mirage Attacks for Highly Accurate Multi-Keyword Query Recovery With Partial Similar Data Against SEabstractSearchable encryption (SE) allows users to perform private queries on encrypted databases. Although SE schemes can protect data privacy, some often pursue high performance while allowing certain leakages, such as search and access patterns. Exploiting such leakage together with other knowledge similar to the user’s database, an attacker can recover queries. State-of-the-art attacks (Nie et al., USENIX’ 24) on single-keyword queries achieve accuracies exceeding 90%. More recently, the community has focused on the more challenging attack of recovering multikeyword queries, with the most advanced attacks (Liu et al., TIFS’ 25) achieving over 80% accuracy. Although these attacks can effectively recover queries, they all rely on a large amount of similar document, requiring the attacker to possess documents equivalent in volume to the database. This naturally raises a question: Can we achieve higher-accuracy attacks using less similar data? Less information makes attacks easier to implement. Motivated by this, we present Mirage, an attack that recovers both single-keyword and multi-keyword queries while requiring only partial similar data. Our core idea is to first identify some special queries and design a series of novel algorithms to recover them. Then, partially reconstruct the database index and recover the remaining queries. Extensive experiments conducted across various real-world datasets demonstrate the effectiveness of our attack. The results show that when the attacker observes 51 time intervals and obtains only 0.5% of similar documents in each interval, Mirage achieves 91.6% and 95.4% accuracy on the Enron and Lucene datasets for single-keyword queries, respectively. For multi-keyword queries, Mirage achieves up to 90.7% and 93.5% recovery accuracy, respectively. Hui Zhu 0001, Songnian Zhang, Yandong Zheng, Mingqin Hou, Wei Xu 0042, Hui Li 0006 |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2025 | Achieving Efficient BGV Scheme with Semi-Honest TEE AssistanceabstractFully Homomorphic Encryption (FHE) allows computations on encrypted data. However, FHE faces significant performance overhead due to the time-consuming homomorphic multiplications and bootstrapping operations, which limits its practical applications. Meanwhile, many FHE optimization schemes have been proposed based on a Trusted Execution Environment (TEE), but they heavily rely on the trustworthiness of the TEE. In outsourcing scenarios, the TEE owner might be a third-party service provider, posing a risk of data leakage for users. To address these issues, we propose an efficient BGV scheme with semi-honest TEE assistance. We consider the TEE to be honest-but-curious, meaning the TEE owner might eavesdrop on the user's data during computation. Based on this assumption, we design a new BGV homomorphic multiplication protocol that operates in a smaller ciphertext space without compromising the original security level and substantially improves computational efficiency. With the assistance of a semi-honest TEE, our multiplication protocol eliminates the need for the switch key, reducing initialization time and supporting an unlimited number of homomorphic multiplication operations without bootstrapping. Furthermore, we provide a security proof using a simulationbased real/ideal world model. Extensive experiments show that our scheme outperforms mainstream libraries such as SEAL, OpenFHE, and Helib. Hui Zhu 0001, Fengwei Wang, Songnian Zhang, Wei Xu 0042, Hui Li 0006 |
ICC | 5 |
| 2025 | SGBoost+: Efficient and Privacy-Preserving Vertical Boosting Trees for Federated Outsourced Training and InferenceabstractVertical federated learning for boosting trees has gained significant attention due to its ability to enable participants to collaboratively train high-quality models while preserving data privacy. However, existing privacy-preserving vertical boosting tree schemes suffer from high computation and communication costs or potential security vulnerabilities. Recently, SGBoost, a federated outsourced training and inference scheme, was proposed to address these challenges. However, its performance and security still require significant improvements. Therefore, we propose SGBoost+, an efficient and privacy-preserving vertical boosting tree framework for federated outsourced training and inference. Building upon the strengths of SGBoost, we introduce an RLWE-based lossless and secure internal node construction and an efficient oblivious inference algorithm to finish the model training and inference, significantly enhancing both security and efficiency. To reduce communication cost, we design a ciphertext compression algorithm for model training, which drastically minimizes data transmission costs. Additionally, we analyze the security of a symmetric encryption scheme, specify the required security conditions and parameters, and optimize our model inference based on its improved and secure version. Detailed security analysis confirms that SGBoost+offers strong privacy guarantees. Extensive experiments demonstrate that SGBoost+achieves efficient model training and inference with significantly lower computation and communication costs compared to state-of-the-art schemes. Wei Xu 0042, Hui Zhu 0001, Jiaqi Zhao 0005, Yandong Zheng, Fengwei Wang, Baishun Sun, Songnian Zhang, Dengguo Feng |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2024 | Efficient and Lossless Integrity-preserving Training Scheme for High-dimensional Logistic Regression over Vertical DataabstractLogistic regression is a widely used and efficient machine learning algorithm for data analysis. However, data is often distributed among different participants, and building high-quality models often involves data interaction, which might easily lead to sensitive information leakage. While some schemes have been proposed to address the data security issue through privacy-preserving technologies such as multi-party computation, homomorphic encryption, and trusted execution environment (TEE), they still suffer from some limitations in terms of security, efficiency, and model integrity and are not suitable for the secure outsourcing training. Therefore, in this paper, we propose an efficient and lossless integrity-preserving training for high-dimensional logistic regression over vertically partitioned data. First, we utilize a block strategy to encrypt the data in batches and outsource them to a cloud with TEE to achieve the non-interactive training. Considering TEE’s limited memory, we employ random sampling technology to achieve mini-batch training within TEE, which ensures the data security. We find that the cloud might launch the active attacks to drop the model accuracy, such as lazy attack and replacement attack. To prevent the attacks, we introduce an integrity protection mechanism to verify data integrity and reduce verification frequency using a mixed-probability checking method, including fixed and random probability checking. Detailed analysis has confirmed that our scheme is secure and integrity-preserving. Extensive experiments have demonstrated that our scheme can achieve lossless, non-interactive, and efficient model training. Wei Xu 0042, Hui Zhu 0001, Ruikang Liu, Yandong Zheng, Fengwei Wang, Dengguo Feng |
GLOBECOM | 1 |
| 2024 | Toward Privacy-Preserving and Verifiable XGBoost Training for Horizontal Federated LearningabstractXGBoost, a widely-used machine learning algorithm, has been applied across various fields. To develop a high-quality model, federated learning is employed, allowing participants to keep their data localized while sharing gradients for global model training. However, recent studies have shown that gradients can leak sensitive information, and even when protected, they might still be vulnerable to certain attacks. While several privacy-preserving and verifiable schemes have been proposed to address this issue, they still face limitations in terms of computational and communication costs in certain scenarios. To solve the problem, we propose a privacy-preserving and verifiable XGBoost training scheme for horizontal federated learning. First, we propose a masking with one-time padding protocol (MOTP) to securely aggregate gradients, supporting both online and offline modes for generating random masks. Both approaches reduce the communication costs and improve the training efficiency. Next, we present a linear aggregation verification protocol (LAVP) to ensure the gradient integrity, which avoids complex computations, thereby improving verification efficiency. Building on MOTP and LAVP, we propose a privacy-preserving and verifiable internal node construction algorithm to train a XGBoost model for horizontal data. Detailed security analysis and extensive experiments demonstrate that our scheme can achieve the privacy-preserving, verifiable, lossless, and efficient XGBoost training for horizontally partitioned data. Wei Xu 0042, Hui Zhu 0001, Fengwei Wang, Dengguo Feng, Hui Li 0006 |
TrustCom | 1 |
| 2024 | Enhancing paillier to fully homomorphic encryption with semi-honest TEE
Yunyi Fang, Hui Zhu 0001, Wei Xu 0042, Yandong Zheng, Xingdong Liu |
Peer Peer Netw. Appl. | 4 |
| 2024 | ToNN: An Oblivious Neural Network Prediction Scheme With Semi-Honest TEEabstractWith the rapid advancements in machine learning and the widespread adoption of Model-as-a-Service (MaaS) platforms, there has been significant attention on convolutional neural network (CNN) inference services. However, traditional inference services over plaintext data and models are susceptible to the risks of data and model leakage. Although several privacy-preserving CNN inference schemes utilizing trusted execution environment (TEE) and cryptography have been proposed, their security models and performance still have limitations in some scenarios. Aiming at the above challenges, we present an oblivious neural network prediction scheme with semi-honest TEE, namely ToNN, which ensures the security of users’ inputs, outputs, and the model itself. Specifically, based on the limited memory of the TEE, we design secure protocols to perform CNN calculations securely and efficiently, which are friendly to support the single instruction multiple data technique. Additionally, we propose a look-up-table method to optimize the convolution and pooling layers calculations. A detailed security analysis under the simulation-based real/ideal worlds model shows that ToNN can achieve the desired security. Extensive simulation results further demonstrate that ToNN can improve the performance of linear calculations by$\textbf {4.86}\times $and non-linear calculation by$\textbf {37.68}\times $, and can be implemented effectively with low computation and communication costs. Wei Xu 0042, Hui Zhu 0001, Yandong Zheng, Fengwei Wang, Jiafeng Hua, Dengguo Feng, Hui Li 0006 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2024 | ELXGB: An Efficient and Privacy-Preserving XGBoost for Vertical Federated LearningabstractWith the rapid growth of Internet data volumes, Big Data analysis technologies have gradually permeated all aspects of life. However, the existence of data silos and the promulgation of relevant regulations make it challenging to apply these technologies. In this context, federated learning provides a feasible solution. Especially, XGBoost schemes for vertical federated learning have attracted much attention due to the widespread use of XGBoost. However, these schemes have limitations in terms of security or efficiency. To address these issues, we propose an efficient and privacy-preserving vertical federated learning framework based on the XGBoost algorithm, namely ELXGB, which achieves secure data alignment, XGboost training, and inference services. First, we design two node split algorithms based on homomorphic encryption and differential privacy, which securely and efficiently achieve tree node generation to construct the global model. Then, we utilize attribute obfuscation and direction obfuscation to achieve a secure inference algorithm, which avoids sensitive information leakage and protects the global model. Additionally, the global model of ELXGB is designed to be centralized, which does not require all participants to stay online for inference. Detailed security analysis demonstrates that ELXGB is privacy-preserving. Moreover, extensive experiments on real-world datasets indicate that ELXGB achieves high efficiency without sacrificing model accuracy. Wei Xu 0042, Hui Zhu 0001, Yandong Zheng, Fengwei Wang, Jiaqi Zhao 0005, Zhe Liu 0001, Hui Li 0006 |
IEEE Trans. Serv. Comput. | 1 |
| 2023 | SGBoost: An Efficient and Privacy-Preserving Vertical Federated Tree Boosting FrameworkabstractAiming at balancing data privacy and availability, Google introduces the concept of federated learning, which can construct global machine learning models over multiple participants while keeping their raw data localized. However, the exchanged parameters in traditional federated learning may still reveal the data information. Meanwhile, the training data are usually partitioned vertically in real-world scenes, which causes difficulties in model construction. To tackle these problems, in this paper, we propose an efficient and privacy-preserving vertical federated tree boosting framework, namely SGBoost, where multiple participants can collaboratively perform model training and query without staying online all the time. Specifically, we first design secure bucket sharing and best split finding algorithms, with which the global tree model can be constructed over vertically partitioned data; meanwhile, the privacy of training data can be well guaranteed. Then, we design an oblivious query algorithm to utilize the trained model without leaking any query data or results. Moreover, SGBoost does not require multi-round interactions between participants, significantly improving the system efficiency. Detailed security analysis shows that SGBoost can well guarantee the privacy of raw data, weights, buckets, and split information. Extensive experiments demonstrate that SGBoost can achieve high accuracy comparable to centralized training and efficient performance. Jiaqi Zhao 0005, Hui Zhu 0001, Wei Xu 0042, Fengwei Wang, Rongxing Lu, Hui Li 0006 |
IEEE Trans. Inf. Forensics Secur. | 3 |