VLDB 2026 Research / reviewers in the wild / expert
Jiaqi Zhao 0005
dblp:27/9676-5
· DBLP profile ↗
20ranked-venue papers
12as first author
20since 2021 · last 2026
0000-0002-1604-1953ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 9 · 4 first-author · 9 since 2021Databases, data management, data science and information retrieval · 5 · 4 first-author · 5 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CoDa: Privacy-preserving multi-dimensional dataset publishing based on consistent data masking
Xiaoyu Kou, Hui Zhu 0001, Jiezhen Tang, Jiaqi Zhao 0005, Fengwei Wang, Hui Li 0006 |
Inf. Sci. | 4 |
| 2026 | Collusion-Resistant Privacy-Preserving Outsourced Training Under Single Cloud With Semi-Honest TEE
Wei Xu 0042, Hui Zhu 0001, Guozhang He, Jingqiang Lin 0001, Jiaqi Zhao 0005, Rongxing Lu, Dengguo Feng |
IEEE Trans. Dependable Secur. Comput. | 5 |
| 2026 | SXGB: Secure and Efficient Vertical Federated XGBoost via Trusted Execution Environments
Jiaqi Zhao 0005, Hui Zhu 0001, Fengwei Wang, Rongxing Lu, Hui Li 0006 |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2026 | Achieving Privacy-Preserving and High-Accuracy Collection of Key-Value Data With Local Differential PrivacyabstractIn the context of the Internet of Things (IoT), the large-scale generation and collection of data can greatly improve the quality of service provided, but they also raise significant concerns about privacy breaches. However, existing privacy-preserving data collection solutions based on local differential privacy (LDP) often struggle to balance security and accuracy when handling composite data types. To address this challenge, in this paper, we propose CSKV, a high-precision and privacy-preserving key-value data collection scheme. Specifically, we first design a padding and sampling protocol to improve data utility. Then, we propose two randomized response mechanisms to safely perturb keys and values in a cohesive and segmented manner. After that, by leveraging the sampling protocol and key-value correlation perturbation, we demonstrate that CSKV can provide secondary privacy amplification. Detailed theoretical analysis verifies the security and effectiveness of CSKV. In addition, extensive performance evaluations are conducted on synthetic and real-world datasets, and the results indicate that our proposed scheme outperforms existing schemes in terms of hit rate and estimation variance. Hui Zhu 0001, Jiaqi Zhao 0005, Mengqian Li, Shuang Zhang 0009, Hui Li 0006 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2026 | PACT: Enhancing Privacy and Efficiency in Tree Evaluation via Secure Parallel Comparison and Oblivious Tree Aggregation
Jiaqi Zhao 0005, Hui Zhu 0001, Fengwei Wang, Yandong Zheng, Hui Li 0006 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2026 | Plog: An Efficient and Privacy-Preserving Collaborative Learning Framework on Vertically Partitioned Graph DataabstractWith the rapid advancement and widespread ap plication of the graph neural network (GNN), the collaborative graph learning, in which multiple parties collaboratively construct a GNN model using their respective graph data, has attracted increasing attention. However, this paradigm also raises significant privacy concerns, as both nodes and edges may contain sensitive personal information, while existing privacy preserving schemes often come at the cost of degraded model performance or substantial system overhead. Therefore, this paper proposes an efficient and privacy-preserving collaborative, and Hui Li, Member, IEEE, Xiaoyu Kou Social Platform learning framework on vertically partitioned graph data, dubbed Plog. Specifically, we first design a decomposition algorithm to split the sparse adjacency matrix into the summation of multiple independent permutations, which are lightweight, parallelizable, and well-suited for secure multi-party computation. Building on this, a weighted oblivious batch permutation protocol is carefully customized based on correlated randomness to securely and efficiently compute adjacency matrix multiplications, addressing the core efficiency bottleneck in GNN inference and training. The selective security of Plog is formally verified under the ideal-real paradigm. Extensive experimental results on three real world datasets demonstrate that compared to the state-of-the art scheme, Plog can reduce online communication rounds by 46% and achieve a 1.73× speedup in the overall inference and training time. Jiaqi Zhao 0005, Hui Zhu 0001, Xiaoyu Kou, Haonan Yan, Fengwei Wang, Hui Li 0006 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2025 | SplitAD: A lightweight and privacy-enhancing vertical federated anomaly detection framework based on hierarchical autoencoders
Jiaqi Zhao 0005, Hui Zhu 0001, Jiezhen Tang, Fengwei Wang, Hui Li 0006 |
Inf. Sci. | 1 |
| 2025 | SGBoost+: Efficient and Privacy-Preserving Vertical Boosting Trees for Federated Outsourced Training and InferenceabstractVertical federated learning for boosting trees has gained significant attention due to its ability to enable participants to collaboratively train high-quality models while preserving data privacy. However, existing privacy-preserving vertical boosting tree schemes suffer from high computation and communication costs or potential security vulnerabilities. Recently, SGBoost, a federated outsourced training and inference scheme, was proposed to address these challenges. However, its performance and security still require significant improvements. Therefore, we propose SGBoost+, an efficient and privacy-preserving vertical boosting tree framework for federated outsourced training and inference. Building upon the strengths of SGBoost, we introduce an RLWE-based lossless and secure internal node construction and an efficient oblivious inference algorithm to finish the model training and inference, significantly enhancing both security and efficiency. To reduce communication cost, we design a ciphertext compression algorithm for model training, which drastically minimizes data transmission costs. Additionally, we analyze the security of a symmetric encryption scheme, specify the required security conditions and parameters, and optimize our model inference based on its improved and secure version. Detailed security analysis confirms that SGBoost+offers strong privacy guarantees. Extensive experiments demonstrate that SGBoost+achieves efficient model training and inference with significantly lower computation and communication costs compared to state-of-the-art schemes. Wei Xu 0042, Hui Zhu 0001, Jiaqi Zhao 0005, Yandong Zheng, Fengwei Wang, Baishun Sun, Songnian Zhang, Dengguo Feng |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2025 | COKV: Key-Value Data Collection With Condensed Local Differential PrivacyabstractLocal differential privacy (LDP) provides lightweight and provable privacy protection and has wide applications in private data collection. Key-value data, as a popular NoSQL structure, requires simultaneous frequency and mean estimations of each key, which poses a challenge to traditional LDP-based collection methods. Despite many schemes proposed for the privacy protection of key-value data, they inadequately solve the condensed perturbation for keys and the advanced combination of privacy budgets, leading to suboptimal estimation accuracy. To address this issue, we propose an efficient key-value collection scheme (COKV) with tight privacy budget composition. In our scheme, we first design a padding and sampling protocol for key-value data to avoid privacy budget splitting. Second, to enhance the utility of key perturbation, we design a key perturbation primitive and optimize the perturbation range to improve computational efficiency. After that, we propose a key-value association perturbation algorithm whose value perturbation strategy guarantees the output expectation equals the original value. Finally, we demonstrate that through a tight privacy budget composition, COKV can provide higher data utility under the same privacy level. Theoretical analysis shows that COKV possesses lower frequency and mean estimations variance. Extensive experiments on both synthetic and real-world datasets also indicate that COKV outperforms the current state-of-the-art methods for secure key-value data collection. Hui Zhu 0001, Jiaqi Zhao 0005, Rongxing Lu, Yandong Zheng, Jiezhen Tang, Hui Li 0006 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2024 | UNIRE: Secure Trajectory-User Linking Model Aggregation with Knowledge TransferabstractMachine learning-based trajectory-user linking (TUL) task, which explores human mobility patterns and identifies known users from anonymous trajectories, is widely applied in location-based personalized recommendation systems. However, TUL models implicitly and inadvertently retain users historical trajectories in training data, which could be revealed under malicious inference and analysis. To address this issue, we propose UNIRE, a knowledge transfer-style TUL model aggregation framework to protect sensitive training data. Specifically, teacher TUL models are trained with private users trajectories and then utilized to collaboratively annotate a sensitive-trajectories-isolated dataset. Leveraging this teacher-labeled dataset, student TUL model can be trained and released as the secure aggregation of teacher models, preventing attackers from arbitrarily accessing or exposing the private trajectories. Meanwhile, differential privacy is applied for safeguarding the knowledge transfer process, ensuring that aggregation information from individual teacher models is indistinguishable. Furthermore, we introduce trajectory Equalization and Pseudo Classification as data alignment and accuracy optimization mechanisms for training and aggregating, respectively. Finally, security analysis and extensive experiments indicate that UNIRE framework achieves effective sensitive training data preservation and performance-improved TUL models aggregation, compared to state-of-the-art baselines. Jiezhen Tang, Hui Zhu 0001, Yandong Zheng, Fengwei Wang, Jiaqi Zhao 0005, Hui Li 0006 |
TrustCom | 6 |
| 2024 | ELXGB: An Efficient and Privacy-Preserving XGBoost for Vertical Federated LearningabstractWith the rapid growth of Internet data volumes, Big Data analysis technologies have gradually permeated all aspects of life. However, the existence of data silos and the promulgation of relevant regulations make it challenging to apply these technologies. In this context, federated learning provides a feasible solution. Especially, XGBoost schemes for vertical federated learning have attracted much attention due to the widespread use of XGBoost. However, these schemes have limitations in terms of security or efficiency. To address these issues, we propose an efficient and privacy-preserving vertical federated learning framework based on the XGBoost algorithm, namely ELXGB, which achieves secure data alignment, XGboost training, and inference services. First, we design two node split algorithms based on homomorphic encryption and differential privacy, which securely and efficiently achieve tree node generation to construct the global model. Then, we utilize attribute obfuscation and direction obfuscation to achieve a secure inference algorithm, which avoids sensitive information leakage and protects the global model. Additionally, the global model of ELXGB is designed to be centralized, which does not require all participants to stay online for inference. Detailed security analysis demonstrates that ELXGB is privacy-preserving. Moreover, extensive experiments on real-world datasets indicate that ELXGB achieves high efficiency without sacrificing model accuracy. Wei Xu 0042, Hui Zhu 0001, Yandong Zheng, Fengwei Wang, Jiaqi Zhao 0005, Zhe Liu 0001, Hui Li 0006 |
IEEE Trans. Serv. Comput. | 5 |
| 2024 | Efficient and Privacy-Preserving Federated Learning Against Poisoning AdversariesabstractThe ever-growing data scale and increasingly strict privacy restraint have recently drawn extensive attention to federated learning (FL) as a multi-party machine learning paradigm for achieving high-quality model construction without data collection. Nevertheless, uploading local models in FL can still be exploited by adversaries to infer participants' sensitive data. Furthermore, it is possible for malicious participants to manipulate the global model by submitting poisonous local models. To tackle these challenges, this paper proposes an efficient and privacy-preserving federated learning framework against poisoning adversaries, namely ELFL, which can ensure the confidentiality of local models while effectively resisting data poisoning attacks. Specifically, we first design a grouped secure aggregation algorithm, through which the aggregation server can compute the summations of local models inside logic groups but cannot see individual ones. Then, based on grouped aggregations, our poisoning defense mechanism could detect and quickly phase out malicious participants from training candidates. Moreover, the computational complexity of participants is independent of their total number, so it is suitable for large-scale scenes. Detailed security analysis demonstrates the security of ELFL. Experimental results show that ELFL could maintain a high accuracy against representative data poisoning attacks, and its computational and communication overhead is indeed low. Jiaqi Zhao 0005, Hui Zhu 0001, Fengwei Wang, Yandong Zheng, Rongxing Lu, Hui Li 0006 |
IEEE Trans. Serv. Comput. | 1 |
| 2023 | Efficient and Privacy-Preserving Logistic Regression Prediction over Vertically Partitioned DataabstractThe explosive growth of data scales has recently drawn extensive attention to machine learning (ML) prediction services for achieving effective and efficient decision-making. However, the potential risk of model or data leakage, as well as vertically partitioned query data in real-world scenes, bring many challenges to ML prediction. Therefore, this paper proposes an efficient and privacy-preserving logistic regression prediction scheme, through which query institutions with vertically partitioned data can enjoy secure, fast, and high-accuracy prediction services from one cloud service provider. Specifically, we take the efficiency advantage of the trusted execution environments (TEE) technique under a semi-honest model, and design a series of masking and recovery methods based on homomorphic encryption to guarantee privacy inside and outside the enclave. Moreover, the ciphertext packing technique is integrated into our scheme for processing predictions parallelly, which further reduces communication and computational overhead. As a result, our scheme will not introduce additional strong security assumptions to TEE; meanwhile, it achieves a trade-off between security and efficiency. Based on the simulation-based paradigm, we formally prove the selective security of our scheme. Extensive experiments on real-world and synthetic datasets show that our prediction accuracy is comparable with plaintext prediction, and our scheme has at least a 30× running speedup compared to the existing representative schemes. Jiaqi Zhao 0005, Hui Zhu 0001, Fengwei Wang, Rongxing Lu, Hui Li 0006 |
GLOBECOM | 1 |
| 2023 | Efficient and privacy-preserving tree-based inference via additive homomorphic encryption
Jiaqi Zhao 0005, Hui Zhu 0001, Fengwei Wang, Rongxing Lu, Hui Li 0006 |
Inf. Sci. | 1 |
| 2023 | VFLR: An Efficient and Privacy-Preserving Vertical Federated Framework for Logistic RegressionabstractWith the explosive growth of data volume and computing capability, federated learning, which involves constructing global models over multiple data islands, has demonstrated its advantages and vast prospects in the field of machine learning. However, due to commonly vertically partitioned data, coupled with privacy concerns about data leakage, there are still some challenging issues in traditional federated learning. To tackle these challenges, in this article, we propose an efficient and privacy-preserving vertical federated learning framework for logistic regression, named VFLR, where multiple participants can collaboratively perform global model training and query over their vertically partitioned data. Specifically, we first design a data aggregation matrix construction algorithm, with which the vertically partitioned data can be aggregated for high-accuracy global model training. Then, by utilizing a novel symmetric homomorphic encryption, our framework can ensure that the whole training and query processes do not leak any private information. Moreover, based on the data aggregation matrix, multi-round interactions are not required in VFLR, improving training efficiency significantly. Detailed security analysis shows that VFLR can well protect data and model information from inference attacks. In addition, extensive experiments demonstrate that VFLR has high training and query accuracy and low computation and communication overhead. Jiaqi Zhao 0005, Hui Zhu 0001, Fengwei Wang, Rongxing Lu, Ermei Wang, Hui Li 0006 |
IEEE Trans. Cloud Comput. | 1 |
| 2023 | SGBoost: An Efficient and Privacy-Preserving Vertical Federated Tree Boosting FrameworkabstractAiming at balancing data privacy and availability, Google introduces the concept of federated learning, which can construct global machine learning models over multiple participants while keeping their raw data localized. However, the exchanged parameters in traditional federated learning may still reveal the data information. Meanwhile, the training data are usually partitioned vertically in real-world scenes, which causes difficulties in model construction. To tackle these problems, in this paper, we propose an efficient and privacy-preserving vertical federated tree boosting framework, namely SGBoost, where multiple participants can collaboratively perform model training and query without staying online all the time. Specifically, we first design secure bucket sharing and best split finding algorithms, with which the global tree model can be constructed over vertically partitioned data; meanwhile, the privacy of training data can be well guaranteed. Then, we design an oblivious query algorithm to utilize the trained model without leaking any query data or results. Moreover, SGBoost does not require multi-round interactions between participants, significantly improving the system efficiency. Detailed security analysis shows that SGBoost can well guarantee the privacy of raw data, weights, buckets, and split information. Extensive experiments demonstrate that SGBoost can achieve high accuracy comparable to centralized training and efficient performance. Jiaqi Zhao 0005, Hui Zhu 0001, Wei Xu 0042, Fengwei Wang, Rongxing Lu, Hui Li 0006 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2022 | ACCEL: an efficient and privacy-preserving federated logistic regression scheme over vertically partitioned data
Jiaqi Zhao 0005, Hui Zhu 0001, Fengwei Wang, Rongxing Lu, Hui Li 0006, Zhongmin Zhou, Haitao Wan |
Sci. China Inf. Sci. | 1 |
| 2022 | An Efficient and Privacy-Preserving Route Matching Scheme for Carpooling ServicesabstractWith the popularity of intelligent terminals and the advances of mobile Internet, carpooling service, which reduces the travel cost of each user by allowing multiple users to share one car, has received considerable attention and makes our life more convenient. However, the vigorous development of carpooling services still faces severe challenges in users’ location or route privacy. In this article, we propose an efficient and privacy-preserving route matching scheme called TAROT for carpooling services. With TAROT, users can enjoy high-quality carpooling services while without revealing sensitive location and route information. Specifically, based on a Goldwasser–Micali-based equality determination algorithm (GMEDA), we design an accurate similarity computation algorithm (ASCA), which allows users to get accurate carpooling results over ciphertexts. Meanwhile, the reverse Minhash (RM) method is also designed to construct a dissimilar route filter algorithm (DRFA), which can filter out dissimilar routes in advance and reduce computational costs and communication overheads. Security analysis shows that TAROT can protect users’ location privacy. In addition, TAROT is also evaluated with many random maps, and the simulation results demonstrate that TAROT is highly efficient. Qi Xu 0002, Hui Zhu 0001, Yandong Zheng, Jiaqi Zhao 0005, Rongxing Lu, Hui Li 0006 |
IEEE Internet Things J. | 4 |
| 2022 | CORK: A privacy-preserving and lossless federated learning scheme for deep neural network
Jiaqi Zhao 0005, Hui Zhu 0001, Fengwei Wang, Rongxing Lu, Hui Li 0006, Jingwei Tu |
Inf. Sci. | 1 |
| 2022 | PVD-FL: A Privacy-Preserving and Verifiable Decentralized Federated Learning FrameworkabstractOver the past years, the increasingly severe data island problem has spawned an emerging distributed deep learning framework—federated learning, in which the global model can be constructed over multiple participants without directly sharing their raw data. Despite its promising prospect, there are still many security challenges in federated learning, such as privacy preservation and integrity verification. Furthermore, federated learning is usually performed with the assistance of a center, which is prone to cause trust worries and communicational bottlenecks. To tackle these challenges, in this paper, we propose a privacy-preserving and verifiable decentralized federated learning framework, named PVD-FL, which can achieve secure deep learning model training under a decentralized architecture. Specifically, we first design an efficient and verifiable cipher-based matrix multiplication (EVCM) algorithm to execute the most basic calculation in deep learning. Then, by employing EVCM, we design a suite of decentralized algorithms to construct the PVD-FL framework, which ensures the confidentiality of both global model and local update and the verification of every training step. Detailed security analysis shows that PVD-FL can well protect privacy against various inference attacks and guarantee training integrity. In addition, the extensive experiments on real-world datasets also demonstrate that PVD-FL can achieve lossless accuracy and practical performance. Jiaqi Zhao 0005, Hui Zhu 0001, Fengwei Wang, Rongxing Lu, Zhe Liu 0001, Hui Li 0006 |
IEEE Trans. Inf. Forensics Secur. | 1 |