Meng Hao 0001

dblp:184/7209-1 · DBLP profile ↗
← Back
45ranked-venue papers
9as first author
42since 2021 · last 2026
0000-0002-1545-1926ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 20 · 5 first-author · 20 since 2021Computer networks · 19 · 2 first-author · 17 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Efficient and Verifiable Data Statistical Analysis via Zero-knowledge Proofs
Hanxiao Chen 0001, Rui Zhang 0086, Pengzhi Xing, Meng Hao 0001, Hongwei Li 0001
ICC4
2026 Efficient Privacy-Preserving Genetic Analysis via Distributed Function Secret Sharing
Shenghao Wu, Pengzhi Xing, Meng Hao 0001, Hanxiao Chen 0001, Wenbo Jiang 0001, Hongwei Li 0001
ICC3
2026 Efficient Fuzzy Private Set Intersection from Secret-Shared OPRF
Xinpeng Yang, Meng Hao 0001, Chenkai Weng, Robert H. Deng, Yonggang Wen 0001, Tianwei Zhang 0004
SP2
2026 An Advanced Gradient Leakage Attack Against Duplicate Labels via Model Outputs Reconstruction
abstract
Federated learning (FL) is a prevalent distributed machine learning framework that allows multiple clients to train one model by uploading gradients without sharing data, enabling cooperative learning while preserving the training data privacy. Nevertheless, recent research has revealed that shared gradients can still expose clients' private training data. These attacks, however, often become ineffective in two practical scenarios: (1) gradients are computed on high-resolution data; (2) labels are duplicated within the attacked batch. In this work, we introduce an advancedGradientLeakageAttack againstDuplicate labels (GLAD), which can effectively recover high-resolution training data from gradients while considering duplicate labels, making it applicable in more realistic FL scenarios. The key technique ofGLADis to formalize the relationships between model outputs, gradients, model parameters, and training data labels. Based on these relationships,GLADfurther reconstructs the model outputs and inverts the reconstructed model outputs back to the corresponding model inputs. Our method can achieve state-of-the-art recovery accuracy while ensuring efficiency. Extensive experimental results demonstrate thatGLADcan reconstruct images of 224× 224pixels with a batch size of 256 with duplicate labels. Our source code is available athttps://github.com/SuperX612/GLAD.
Kunlan Xiang, Haomiao Yang, Meng Hao 0001, Zikang Ding, Hongwei Li 0001, Qingchuan Zhao, Tianwei Zhang 0004
IEEE Trans. Dependable Secur. Comput.3
2026 HyperSiniel: Guaranteed Output Delivery Comes (Almost) Free in Private Delegation of zkSNARKs
abstract
Zero-knowledge Succinct Non-interactive Argument of Knowledge (zkSNARK) is a powerful cryptographic primitive that enables a prover to convince a verifier that something is true without leaking the private witness. Current zkSNARKs face significant computational costs in generating proofs, which restricts their use in areas like private payments, confidential smart contracts, and anonymous credentials. Private delegation offers a practical solution by outsourcing the heavy computation to powerful external workers without leaking any private information. In this work, we propose HyperSiniel, an efficient private delegation framework for general zkSNARKs that achieves a new feature called guaranteed output delivery (GOD). HyperSiniel is designed to be compatible with any universal zkSNARKs constructed from a polynomial interactive oracle proof (PIOP) and a polynomial commitment scheme (PCS). It enables a computationally limited delegator to outsource proof generation to several workers in a fully non-interactive and privacy-preserving manner. Compared to the most state-of-the-art frameworks (e.g., Siniel [NDSS'25]), HyperSiniel ensures that the delegator always receives a correct proof, regardless of malicious worker behavior. We implement HyperSiniel and compare the performance with Siniel across varying bandwidths and circuit sizes. Under low-bandwidth conditions (10MBps), HyperSiniel incurs only an additional 25% overhead compared with Siniel, while the total running time of HyperSiniel is almost identical to Siniel under high-bandwidth settings (1000MBps). These results show that the strong robustness guarantee of GOD in HyperSiniel comes almost for free, making it a practical and secure solution for real-world zkSNARK delegation.
Yunbo Yang, Yuejia Cheng, Junkai Liang, Kailun Wang, Xuanming Liu, Xiaoguo Li, Jianfei Sun, Xiaolei Dong, Zhenfu Cao, Meng Hao 0001, Guomin Yang, Robert H. Deng, Kui Ren 0001
IEEE Trans. Dependable Secur. Comput.11
2026 Sanitizer: Blazing-Fast, Private, and Robust Federated Learning
abstract
Recently, private and robust federated learning (FL) schemes have been proposed to address privacy inference and Byzantine attacks simultaneously. However, existing schemes are inefficient in private and robust aggregation protocols due to the employment of heavy cryptographic techniques. To approach the above problem, we propose Sanitizer, an efficient, private, and robust FL framework. Specifically, we first design a Byzantine-robust defense for communication-efficient sign-based FL. We further propose a customized private and robust aggregation scheme built on our Byzantine-robust defense for FL. The core of our construction is two new efficient protocols, i.e.,high-dimensional boolean summationandweighted boolean majority vote, which serve as the main building blocks of Sanitizer. Extensive evaluations on real-world datasets demonstrate that Sanitizer is blazing fast, achieving 19 ∼ 23× less runtime compared to the state-of-the-art. Meanwhile, Sanitizer achieves the same accuracy as the plaintext and superior Byzantine robustness against various classic attacks.
Hanxiao Chen 0001, Hongwei Li 0001, Meng Hao 0001, Jia Hu 0004, Hao Ren 0001, Haomiao Yang, Tianwei Zhang 0004, Guowen Xu
IEEE Trans. Inf. Forensics Secur.3
2026 Conan: Secure and Reliable Machine Learning Inference Against Malicious Service Providers
abstract
In the Machine Learning as a Service paradigm, a service provider (e.g., a server) hosting a model offers inference APIs to clients, who can send their queries and receive the inference results. While most recent secure inference works focus on addressing privacy issues, they overlook the importance of checking the service quality and reliability. A malicious server may deviate from the protocol specification to deliberately provide incorrect services such as using low-quality models. Thus, it is necessary to design new solutions to empower clients to verify the server’s model accuracy and inference integrity while protecting both parties’ privacy. We present Conan, a new secure and reliable inference framework against malicious servers to achieve accuracy verification, inference integrity, and privacy simultaneously. In Conan, the server first commits to the model and proves in zero-knowledge that the committed model achieves the claimed accuracy. Then both parties perform secure inference on the committed model against the malicious server. To instantiate the above framework, we design generic maliciously secure two-party computation (2PC) protocols with a fixed corrupted party, which may be of independent interest. Our protocols achieve high efficiency by utilizing the advantage that the semi-honest party can check the behavior of the corrupted party. Furthermore, they support both arithmetic and Boolean circuit evaluation, a crucial attribute for secure inference on complicated machine learning models. We implement the fixed-corruption 2PC protocols for our secure and reliable inference. The experimental results show 1 ~ 2 orders of magnitude improvements over conventional maliciously secure protocols in terms of communication and computation costs.
Hanxiao Chen 0001, Hongwei Li 0001, Meng Hao 0001, Pengzhi Xing, Jia Hu 0004, Wenbo Jiang 0001, Tianwei Zhang 0004, Guowen Xu
IEEE Trans. Inf. Forensics Secur.3
2026 T3AT: Threshold-Authorized, Threshold-Redeemable, and Non-Transferable Anonymous Tokens
abstract
Anonymous authentication mechanisms play an increasingly critical role in digital ecosystems by enabling users to prove eligibility without revealing identity information. Anonymous tokens serve as fundamental cryptographic primitives for privacy-preserving access control. However, existing solutions often rely on trusted hardware or suffer from centralization issues, such as single points of failure (SPoF) and strong trust assumptions, to enforce non-transferability. In this work, we construct a threshold BBS+ signing protocol using verifiable multiplication-to-addition (MtA) techniques derived from vector oblivious linear evaluation (VOLE). The security of the proposed threshold signature scheme is rigorously established within the Universal Composability (UC) framework. Building upon this foundation, we introduce the first threshold-authorized and threshold-redeemable, and non-transferable anonymous tokens named T3AT. T3AT enables collaborative issuance and verification in the malicious adversary and dishonest majority settings while achieving non-transferability, unlinkability, and unforgeability without relying on trusted hardware or centralized authorities. Our performance evaluation demonstrates the practicality, efficiency, and scalability of T3AT, effectively bridging the gap between anonymous tokens and threshold-based authorization and authentication for privacy-enhanced access control.
Jian Shen 0001, Jianting Ning, Meng Hao 0001, Leo Yu Zhang
IEEE Trans. Inf. Forensics Secur.4
2025 SecInfer: Secure and Efficient Model Inference on Vertically Partitioned Data
abstract
Deep learning models have achieved unprecedented success in various domains, such as healthcare and finance. However, deploying model inference in real-world applications, where data is distributed among multiple entities, poses significant privacy concerns. Existing secure model inference work has limitations in computational overhead and scalability, especially when dealing with complex models and multiple parties with vertically partitioned data. In this work, we design and implement an efficient and scalable secure inference framework for vertically partitioned data, supporting execution with a large number of parties. Our work considers a semi-honest setting with all-but-one corruptions. The core of our framework is a series of secure and efficient protocols for complex non-linear functions of the model inference, such as ReLU and Maxpool. These protocols are designed based on secure multi-party computation preliminaries, significantly enhancing efficiency while maintaining rigorous security guarantees. We conduct comprehensive experiments to evaluate the performance of our framework. Experimental results show that SecInfer substantially improves the communication and computation performance of secure naive inference works by up to 3.71 × and 3.42 ×, respectively.
Robert H. Deng, Hongwei Li 0001, Hanxiao Chen 0001, Meng Hao 0001, Pengzhi Xing, Jia Hu 0004, Rui Zhang 0086, Wenbo Jiang 0001
ICC4
2025 Distributed Function Secret Sharing and Applications
Pengzhi Xing, Hongwei Li 0001, Meng Hao 0001, Hanxiao Chen 0001, Jia Hu 0004
NDSS3
2025 Practical Keyword Private Information Retrieval from Key-to-Index Mappings
Meng Hao 0001, Liqiang Peng, Pengfei Wu 0003, Lei Zhang 0006, Hongwei Li 0001, Robert H. Deng
USENIX Security Symposium1
2025 The Gradient Puppeteer: Adversarial Domination in Gradient Leakage Attacks Through Model Poisoning
Kunlan Xiang, Haomiao Yang, Meng Hao 0001, Shaofeng Li 0001, Haoxin Wang 0004, Zikang Ding, Wenbo Jiang 0001, Tianwei Zhang 0004
IEEE Trans. Inf. Forensics Secur.3
2025 Charge Your Clients: Payable Secure Computation and Its Applications
abstract
The online realm has witnessed a surge in the buying and selling of data, prompting the emergence of dedicated data marketplaces. These platforms cater to servers (sellers), enabling them to set prices for access to their data, and clients (buyers), who can subsequently purchase these data, thereby streamlining and facilitating such transactions. However, the current data market is primarily confronted with the following issues. Firstly, they fail to protect client privacy, presupposing that clients submit their queries in plaintext. Secondly, these models are susceptible to being impacted by malicious client behavior, for example, enabling clients to potentially engage in arbitrage activities. To address the aforementioned issues, we propose payable secure computation, a novel secure computation paradigm specifically designed for data pricing scenarios. It grants the server the ability to securely procure essential pricing information while protecting the privacy of client queries. Additionally, it fortifies the server’s privacy against potential malicious client activities. As specific applications, we have devised customized payable protocols for two distinct secure computation scenarios: Keyword Private Information Retrieval (KPIR) and Private Set Intersection (PSI). We implement our two payable protocols and compare them with the state-of-the-art related protocols that do not support pricing as a baseline. Since our payable protocols are more powerful in the data pricing setting, the experiment results show that they do not introduce much overhead over the baseline protocols. Our payable KPIR achieves the same online cost as baseline, while the setup is about 1.3−1.6× slower than it. Our payable PSI needs about 2× more communication cost than that of baseline protocol, while the runtime is 1.5−3.2× slower than it depending on the network setting.
Liqiang Peng, Meng Hao 0001, Lei Zhang 0006, Dongdai Lin
IEEE Trans. Inf. Forensics Secur.5
2024 Unbalanced Private Set Union with Reduced Computation and Communication
abstract
Private set union (PSU) is a cryptographic protocol that allows two parties to compute the union of their sets without revealing anything else. Despite some efficient PSU protocols that have been proposed, they mainly focus on the balanced setting, where the sets held by the parties are of similar size. Recently, Tu et al. (CCS 2023) proposed the first unbalanced PSU protocol which achieves sublinear communication complexity in the size of the larger set.
Yu Chen 0003, Liqiang Peng, Meng Hao 0001, Anyu Wang 0001, Xiaoyun Wang 0001
CCS5
2024 Benchmark GELU in Secure Multi-Party Computation
abstract
Recently, several technology companies have released online inference services for clients based on Transformer-based large language models, which show excellent performance in various tasks. However, in these services, the inputs usually involve clients’ sensitive information. To address this problem, many works have proposed secure inference on language models such as GPT. For language models, complex mathematical functions like Gaussian Error Linear Unit (GELU) are used extensively and dominate the main cost of secure inference. In this work, we systematically study the existing secure GELU protocols and classify previous methods into two categories: polynomial-based protocols and lookup table (LUT)-based protocols. We point out several important characteristics and tradeoffs for these two classes of secure GELU protocols. Based on these observations and analysis, we propose a new secure GELU protocol, called Simple. The main technique that Simple uses involves a LUT of small size to retrieve approximate polynomials for fitting residual error functions caused by a crude approximation for GELU, which achieves state-of-the-art (SOTA) overhead and accuracy performance. We conduct extensive experiments and benchmark the previous 6 secure GELU protocols. The experimental comparison shows that our Simple protocol achieves 1.1 ∼ 8784.3× computation and 1.4 ∼ 188.8× communication improvements while reducing 1.2∼80.2× errors.
Rui Zhang 0090, Hongwei Li 0001, Meng Hao 0001, Hanxiao Chen 0001, Yuan Zhang 0006, Dianhua Tang
GLOBECOM3
2024 Scalable Zero-knowledge Proofs for Non-linear Functions in Machine Learning
Meng Hao 0001, Hanxiao Chen 0001, Hongwei Li 0001, Chenkai Weng, Yuan Zhang 0006, Haomiao Yang, Tianwei Zhang 0004
USENIX Security Symposium1
2024 Unbalanced Circuit-PSI from Oblivious Key-Value Retrieval
Meng Hao 0001, Liqiang Peng, Hongwei Li 0001, Hanxiao Chen 0001, Tianwei Zhang 0004
USENIX Security Symposium1
2024 SVFGNN: A privacy-preserving vertical federated graph neural network model training framework based on split learning
Yanjun Liu 0010, Hongwei Li 0001, Meng Hao 0001
Peer Peer Netw. Appl.3
2024 Decentralized Multi-Client Functional Encryption for Inner Product With Applications to Federated Learning
abstract
Decentralized multi-client functional encryption for inner product (DMCFE-IP) enables efficient joint functional computation of private inputs in a secure manner without a trusted third party, which has found successful applications, including distributed statistical analysis and machine learning. However, existing DMCFE-IP schemes suffer several drawbacks, such as lack of support for client dropout, requiring cross-client communication for key generation, and poor efficiency and scalability. To address these issues, we propose an efficient and scalable DMCFE-IP, which supports client dropout and non-interactive decentralized partial decryption key generation. Our scheme mainly exploits appropriate underlying cryptographic primitives, including multi-client functional encryption, digital signature, key agreement, secret sharing, and symmetric encryption, with careful integration to achieve the aforementioned two functionalities. We then extend this scheme to enable privacy-preserving federated learning (PPFL) for the cross-silo scenrio. We provide formal security proof for our scheme and evaluate our DMCFE-IP-based PPFL on several real-world datasets. Compared with the state-of-the-art methods, our approach achieves a speedup of 6.12$\sim 43.36\times$in running time.
Xinyuan Qian 0002, Hongwei Li 0001, Meng Hao 0001, Guowen Xu, Haoyong Wang, Yuguang Fang
IEEE Trans. Dependable Secur. Comput.3
2024 SecBNN: Efficient Secure Inference on Binary Neural Networks
abstract
This work studies secure inference on Binary Neural Networks (BNNs), which have binary weights and activations as a desirable feature. Although previous works have developed secure methodologies for BNNs, they still have performance limitations and significant gaps in efficiency when applied in practice. We present SecBNN, an efficient secure two-party inference framework on BNNs. SecBNN exploits appropriate underlying primitives and contributes efficient protocols for the non-linear and linear layers of BNNs. Specifically, for non-linear layers, we introduce a secure sign protocol with an innovative adder logic and customized evaluation algorithms. For linear layers, we propose a new binary matrix multiplication protocol, where a divide-and-conquer strategy is provided to recursively break down the matrix multiplication problem into multiple sub-problems. Building on top of these efficient ingredients, we implement and evaluate SecBNN over two real-world datasets and various model architectures under LAN and WAN. Experimental results show that SecBNN substantially improves the communication and computation performance of existing secure BNN inference works by up to$29 \times $and$14 \times $, respectively.
Hanxiao Chen 0001, Hongwei Li 0001, Meng Hao 0001, Jia Hu 0004, Guowen Xu, Tianwei Zhang 0004
IEEE Trans. Inf. Forensics Secur.3
2024 Stealthy Targeted Backdoor Attacks Against Image Captioning
abstract
In recent years, there is an explosive growth in multimodal learning. Image captioning, a classical multimodal task, has demonstrated promising applications and attracted extensive research attention. However, recent studies have shown that image caption models are vulnerable to some security threats such as backdoor attacks. Existing backdoor attacks against image captioning typically pair a trigger either with a predefined sentence or a single word as the targeted output, yet they are unrelated to the image content, making them easily noticeable as anomalies by humans. In this paper, we present a novel method to craft targeted backdoor attacks against image caption models, which are designed to be stealthier than prior attacks. Specifically, our method first learns a special trigger by leveraging universal perturbation techniques for object detection, then places the learned trigger in the center of some specific source object and modifies the corresponding object name in the output caption to a predefined target name. During the prediction phase, the caption produced by the backdoored model for input images with the trigger can accurately convey the semantic information of the rest of the whole image, while incorrectly recognizing the source object as the predefined target. Extensive experiments demonstrate that our approach can achieve a high attack success rate while having a negligible impact on model clean performance. In addition, we show our method is stealthy in that the produced backdoor samples are indistinguishable from clean samples in both image and text domains, which can successfully bypass existing backdoor defenses, highlighting the need for better defensive mechanisms against such stealthy backdoor attacks.
Wenshu Fan, Hongwei Li 0001, Wenbo Jiang 0001, Meng Hao 0001, Shui Yu 0001, Xiao Zhang 0016
IEEE Trans. Inf. Forensics Secur.4
2024 Vertical Federated Learning Across Heterogeneous Regions for Industry 4.0
abstract
This work investigates fine-grained data distribution in real-world federated learning (FL) applications, wherein training samples are distributed across multiple regions, and different clients within each region possess distinct features of local training samples. Furthermore, the datasets and models in these regions often exhibit heterogeneity, characterized by varying label distributions and model architectures, posing challenges to the model construction process. In this article, we propose a vertical federated learning (VFL) framework, named HeteroVFL, to address the data distribution complexities and overcome the hurdles posed by heterogeneous regions. Besides, we enhance the privacy of HeteroVFL by adopting differential privacy, a privacy-preserving technology by injecting measured noise into data based on a stochastic framework. We compare our HeteroVFL with existing solutions on three real-world datasets in simulations. The results demonstrate that HeteroVFL can achieve over 96% accuracy on MNIST, surpassing the accuracy of 90% in the state-of-the-art VFL benchmarks.
Rui Zhang 0086, Hongwei Li 0001, Luoding Tian, Meng Hao 0001, Yuan Zhang 0006
IEEE Trans. Ind. Informatics4
2024 Efficient and Privacy-Preserving Outsourcing of Gradient Boosting Decision Tree Inference
abstract
Recently, outsourcing machine learning inference services to the cloud has become increasingly popular. The inference process, however, remains an open question onhow to effectively protect the model owner's proprietary model, the user's sensitive data, and prediction results. In this work, we propose an efficient and comprehensive privacy-preserving framework for outsourcing Gradient Boosting Decision Tree (GBDT) inference utilizing pseudorandom function and additively homomorphic encryption. Specifically, we first design a transformation method for GBDT to protect the node and structure privacy of the owner's model. On top of the protected model, we further propose customized comparison and random trees permutation protocols, which substantially boost the computation and reduce the communication cost of the outsourcing inference, while preventing the user from inferring privacy associated with GBDT. Besides, we provide rigorous security analysis, and extensive experiments on 7 real-world datasets and various models demonstrating that our scheme achieves up to 36 times less runtime and 69 times less communication compared to the state-of-the-arts.
Shuai Yuan 0009, Hongwei Li 0001, Xinyuan Qian 0002, Meng Hao 0001, Yixiao Zhai, Guowen Xu
IEEE Trans. Serv. Comput.4
2023 Privacy-Preserving and Verifiable Outsourcing Inference Against Malicious Servers
abstract
Outsourcing inference enables users to outsource neural network inference tasks to a service provider (e.g., a remote server). This paradigm has brought enormous convenience and effectively solved the resource-limited issues of users, especially for mobile devices. However, it still suffers from two challenges: (1) The user's input data and inference results contain a large amount of private information, which should not be disclosed. (2) The server in this setting may be malicious and hence violate the inference procedure. For example, the server may use a low-quality model to reduce costs or return wrong inference results. While several privacy-preserving inference works have been proposed, they cannot solve the above two problems at the same time. In this work, we propose PPVI, a secure and verifiable outsourcing inference scheme against malicious service providers. PPVI designs a hybrid check technique for inference integrity verification and employs leveled homomorphic encryption to protect users' privacy. These ingredients together make it possible to protect users' privacy and verify the inference correctness in outsourcing inference simultaneously. Extensive experiment results demonstrate our scheme has an excellent performance in terms of verification accuracy and communication and computational overhead.
Yiyao Liu, Hongwei Li 0001, Meng Hao 0001, Guiqiang Hu
GLOBECOM3
2023 Practical and Privacy-Preserving Density-Based Clustering via Shuffling
abstract
Density-Based Spatial Clustering of Applications with Noise (DBSCAN) is a commonly used density-based clustering algorithm, and the study of its privacy-preserving methods is of practical importance. However, prior works either leak important intermediate results or suffer from intolerable overhead, which makes it difficult to deploy in real-world scenarios. To address this problem, we propose Private-DBSCAN, a practical secure two-party framework for DBSCAN. Specifically, (i) we design an efficient secure comparison protocol for the calculation of the adjacency matrix, which reduces the online communication to only one round and (ii) we employ a secret-shared shuffle protocol to anonymize the data records, which can hide the position relation of elements while avoiding redundant computations. These ingredients allow Private-DBSCAN to achieve practical efficiency and rigorous security at the same time. We implement our protocol and conduct extensive experiments on five datasets, which show that it achieves a$90\sim 340\times$speedup on LAN and$13\sim 73\times$speedup on WAN compared to the state-of-the-art work.
Yingzhe Wang, Hongwei Li 0001, Hanxiao Chen 0001, Meng Hao 0001
GLOBECOM5
2023 Membership Inference Attacks Against the Graph Classification
abstract
Recently, there has been increasing interest in extending deep learning approaches to graph data. Graph representation learning has become an important way to fully utilize the information contained in graph data. Graph Neural Networks (GNNs) have demonstrated significant efficacy in various fields. Previous studies have shown that traditional machine learning models may lead to disclosure of private data, but the privacy risks of GNNs have not received enough attention. In this paper, we propose two attack methods based on the ground-truth label. Our attack approach covers two mainstream attack patterns, including training attack models based on neural networks and setting thresholds. To improve the effectiveness of attacks, we consider incorporating label information of the samples. Since the samples' distributions of different classes are different, the possibility of privacy leakage cannot be treated equally. In neural network-based attacks, we concatenate the label information into the input vector of the attack model. In threshold-based attacks, we set separate thresholds for each label category. We systematically evaluate the performance of membership inference attacks against graph-level classification. Our evaluation on three GNN structures and four benchmark datasets shows that GNNs for graph classification are more vulnerable to the improved attacks. On the DD dataset, our attack achieved an accuracy of 79%. Furthermore, we proposed two defense mechanisms to mitigate the privacy leakage caused by membership inference attacks.
Junze Yang, Hongwei Li 0001, Wenshu Fan, Meng Hao 0001
GLOBECOM5
2023 SecMath: An Efficient 2-Party Cryptographic Framework for Math Functions
abstract
Complex math functions, such as exponential and tanh, are widely applied in machine learning inference tasks like recurrent neural networks (RNNs). Even though a few works have provided secure implementations of these functions, they still suffer from serious performance bottlenecks, leaving efficiency gaps in practice. To approach this issue, we propose SecMath, an efficient 2-party cryptographic framework for complex math functions. Specifically, SecMath contributes novel communication-efficient protocols for secure exponential, sigmoid and tanh operations. These protocols utilize an advanced underlying primitive, silent oblivious transfer, and employ customized optimizations including lookup table techniques to further improve performance. Extensive evaluations show that our new constructions outperform the counterparts in SIRNN (IEEE S&P'21) by a large margin in terms of both communication and computation overhead. For example, the sigmoid operation of SecMath costs 4.15KB communication and less than 0.2 millisecond, which improves SIRNN up to 7.6× in communication and 2.4× in runtime.
Jia Hu 0004, Hongwei Li 0001, Hanxiao Chen 0001, Meng Hao 0001
ICC4
2023 TriFSS: Secure Trigonometric Function Evaluation via Function Secret Sharing
abstract
Trigonometric functions are crucial non-linear operations used in scientific computation and complex machine learning models. However, existing secure computing frameworks either lack support for these operations, or suffer from undesirable performance bottleneck. In this paper, we present an efficient and precise fixed-point framework called TriFSS for securely evaluating trigonometric functions. Specifically, we first design new building blocks based on advanced Function Secret Sharing techniques, achieving reduced communication and computation overhead. Second, with these efficient components, we propose a general evaluation process for these functions, in which periodic properties are fully exploited for better performance. Moreover, we implement the TriFSS framework and conduct extensive experiments. The experimental results show that our protocols achieve at least 23x less communication overhead and 2.8x less latency than the state-of-the-art frameworks, while only resulting in 1 ULP error, which is comparable to floating-point based works.
Pengzhi Xing, Hongwei Li 0001, Meng Hao 0001, Hanxiao Chen 0001, Shengke Zeng
ICC3
2023 Toward Efficient and End-to-End Privacy-Preserving Distributed Gradient Boosting Decision Trees
abstract
Gradient Boosting Decision Trees (GBDTs) are popular machine learning models due to its simplicity, effectiveness, and interpretability. Recently, to alleviate serious privacy leakages in conventional centralized methods, researchers have proposed several privacy-preserving distributed GBDT solutions. However, those approaches still suffer from either insufficient privacy protection or significant runtime and communication overhead. In this paper, we propose an efficient and end-to-end privacy-preserving distributed GBDT framework, called PPD-GBDT, which uses differential privacy, polynomial approximation, and fully homomorphic encryption to achieve comprehensive privacy protection. Specifically, during the boosting phase, we design a novel model preparation method to improve the efficiency of prediction with acceptably slight accuracy/RMSE loss while preventing data owners' corruption. On the other hand, for the prediction phase, we propose a customized secure prediction method, which effectively prevents the malicious server from stealing private information. Besides, we conduct extensive experiments on six datasets and compare with three prior schemes. Evaluation results show that our privacy-preserving scheme achieves lower runtime and up to 40× less communication overhead compared to the state-of-the-arts.
Shuai Yuan 0009, Hongwei Li 0001, Xinyuan Qian 0002, Meng Hao 0001, Yixiao Zhai
ICC4
2023 GuardHFL: Privacy Guardian for Heterogeneous Federated Learning
abstract
Heterogeneous federated learning (HFL) enables clients with different computation and communication capabilities to collaboratively train their own customized models via a query-response paradigm on auxiliary datasets. However, such a paradigm raises serious privacy concerns due to the leakage of highly sensitive query samples and response predictions. We put forth GuardHFL, the first-of-its-kind efficient and privacy-preserving HFL framework. GuardHFL is equipped with a novel HFL-friendly secure querying scheme built on lightweight secret sharing and symmetric-key techniques. The core of GuardHFL is two customized multiplication and comparison protocols, which substantially boost the execution efficiency. Extensive evaluations demonstrate that GuardHFL significantly outperforms the alternative instantiations based on existing state-of-the-art techniques in both runtime and communication cost.
Hanxiao Chen 0001, Meng Hao 0001, Hongwei Li 0001, Kangjie Chen, Guowen Xu, Tianwei Zhang 0004
ICML2
2023 ESA-FedGNN: Efficient secure aggregation for federated graph neural networks
Yanjun Liu 0010, Hongwei Li 0001, Xinyuan Qian 0002, Meng Hao 0001
Peer Peer Netw. Appl.4
2023 PriVDT: An Efficient Two-Party Cryptographic Framework for Vertical Decision Trees
abstract
Privacy-preserving decision trees (DTs) in vertical federated learning are one of the most effective tools to facilitate various privacy-critical applications in reality. However, the main bottleneck of current solutions is their huge overhead, mainly due to the adoption of communication-heavy bit decomposition to realize complex non-linear operations, such as comparison and division. In this paper, we presentPriVDT, an efficient two-party framework for private vertical DT training and inference in the offline/online paradigm. Specifically, we customize several cryptographic building blocks based on an advanced primitive, Function Secret Sharing (FSS). First, we construct an optimized comparison protocol to improve the efficiency via reducing the invocation of FSS evaluations. Second, we devise an efficient and privacy-enhanced division protocol without revealing the range of divisors, which utilizes the above comparison protocol and more importantly new designed FSS-based secure range and digital decomposition protocols. Besides, we further reduce the overhead of linear operations by employing lightweight pseudorandom function-based Beaver’s triple techniques. Building on the above efficient components, we implement thePriVDTframework and evaluate it on 5 real-world datasets on both LAN and WAN. Experimental results show that the end-to-end runtime ofPriVDToutperforms the prior art by$42 \sim 510\times $on LAN and$16 \sim 70\times $on WAN. Moreover,PriVDTprovides comparable accuracy to the non-private setting.
Hanxiao Chen 0001, Hongwei Li 0001, Yingzhe Wang, Meng Hao 0001, Guowen Xu, Tianwei Zhang 0004
IEEE Trans. Inf. Forensics Secur.4
2023 FastSecNet: An Efficient Cryptographic Framework for Private Neural Network Inference
abstract
Private neural network inference has demonstrated great importance in various privacy-critical scenarios. However, the primary challenge remaining in prior works is that the evaluation on encrypted data levies prohibitively high run-time and communication overhead. In this work, we present FastSecNet, an efficient two-party cryptographic framework for private inference in the dealer-based pre-processing setting. Specifically, (1) FastSecNet provides an efficient ReLU protocol for the evalution of non-linear layers, which is built up on a recent advanced cryptographic primitive, function secret sharing (FSS). The core of this construction are an optimized ReLU representation and a customized FSS-based ReLU protocol. (2) For linear layer evaluation, we first propose an efficient PRG-based preprocessing protocol based on the fact that one of the inputs is uniformly random in the offline phase. Then, the online phase only communicates one element and consists of lightweight secret-sharing operations in a ring. Extensive evaluations conducted on 4 real-world datasets and 9 neural network models demonstrate that during the online phase, FastSecNet achieves 14× less runtime and 18× less communication cost compared to the state-of-the-art.
Meng Hao 0001, Hongwei Li 0001, Hanxiao Chen 0001, Pengzhi Xing, Tianwei Zhang 0004
IEEE Trans. Inf. Forensics Secur.1
2022 Fast Secure Aggregation for Privacy-Preserving Federated Learning
abstract
Federated learning (FL) is a new distributed learning paradigm, in which the clients cooperate to conduct the global model without exposing local private data. However, existing privacy inference attacks on FL show that adversaries can still reverse the training data from the submitted model updates. Recently, secure aggregation has been proposed and integrated into the FL framework, which effectively guarantees privacy through various cryptographic techniques, unfortunately at the cost of a large amount of communication and computation. In this paper, we propose a highly efficient secure aggregation scheme, Fast-Aggregate, which significantly reduces the communication and computation overhead while ensuring data privacy and robustness against clients' dropout. Firstly, Fast-Aggregate employs a multi-group regular graph for efficient secure aggregation to boost data parallelism. Secondly, we leverage polynomial multi-point evaluation and fast Lagrange interpolation methods to handle clients' dropout as well as reduce computational complexity. Finally, we adopt an additive mask to guarantee clients' privacy. Riding on the capabilities of Fast-Aggregate, we achieve the secure aggregation overhead of O (N log2$N$), as opposed to O (N2) in the state-of-the-art works. Besides, Fast-Aggregate improves training speed without loss of model quality and provides flexibility to deal with client corruption at the same time.
Yanjun Liu 0010, Xinyuan Qian 0002, Hongwei Li 0001, Meng Hao 0001, Song Guo 0001
GLOBECOM4
2022 CryptoFE: Practical and Privacy-Preserving Federated Learning via Functional Encryption
abstract
Cloud-based services for federated learning has received widespread attention for its ability to collaboratively train a model without collecting users' local data. Although there are existing methods such as homomorphic encryption and secure multi-party computation to address the privacy issues associated with the model parameter exchanging during aggregation, these methods will inevitably lead to huge communication overheads or slow down the training time. Functional encryption (FE) is considered as a new approach to address privacy-preserving federated learning probelms, but the only known FE solution has severe security issues such as leaking master private key, and is impractical. Thus, in this paper, we propose CryptoFE, a cloud-based privacy-preserving federated learning aggregation scheme based on FE. Compared with the only existing FE solution, CryptoFE is efficient in aggregation phase, especially when a high model precision is required, and provides formal privacy guarantees for users' gradients. The experiments with real-world data demonstrate the efficeint performance of our proposed scheme.
Xinyuan Qian 0002, Hongwei Li 0001, Meng Hao 0001, Shuai Yuan 0009, Song Guo 0001
GLOBECOM3
2022 Efficient and Privacy-Preserving Federated Learning with Irregular Users
abstract
Federated learning (FL) enables multiple users to learn a global predictive model by exchanging local updates without disclosing their private datasets. To further protect local updates, several privacy-preserving schemes are proposed and applied in FL. However, a fundamental issue is that irregular users in FL holding low quality updates could decrease the convergence rate, and even worse, damage the model’s usability. While a few works recently explore unified solutions to mitigate the issues of privacy and irregular users meanwhile, the existing methods are still insufficient in terms of accuracy and efficiency. The reasons are two major limitations: inefficiency caused by complex cryptographic algorithms and poor model usability due to ineffective removing strategies for irregular users. To approach the above problems, we propose SAP-IU, a new and efficient federated learning scheme, which achieves irregular users removing and privacy protection at the same time. Specifically, we first design a novel removing algorithm for irregular users called TrustIUthat calculates the weight of each user via the cosine metric. This ensures that the global model is mainly derived from the contributions of high-quality data. We further devise a secure weighted aggregation protocol for TrustIUto protect users’ sensitive information including local updates and data quality. Besides, our scheme is robust to users dropping out during the whole training process. Moreover, extensive experiments show that SAP-IU has a better performance than prior works in terms of training accuracy and efficiency.
Jieyu Xu, Hongwei Li 0001, Meng Hao 0001
ICC4
2022 Secure Feature Selection for Vertical Federated Learning in eHealth Systems
abstract
Privacy-preserving vertical federated learning (VFL) has been widely applied in electronic health (eHealth) systems. However, existing VFL schemes rarely consider the data pre-processing step including feature selection, which will lead to poor convergence rate and even damaging the model utility. In this paper, we propose an efficient and privacy-preserving feature selection scheme for VFL. Specifically, we first propose a general Gini-impurity based feature selection framework, which is compatible with most existing machine learning models in VFL. With the framework, we present two concrete protocols (dubbed πSS−FSand πH−FS, respectively) customized for different eHealth scenarios. πSS−FSexploits a lightweight additive secret sharing technique, such that it can be executed in comparable time as the evaluation of the plaintext scheme. πH−FSis a hybrid feature selection protocol that additionally utilizes a linear homomorphic encryption technique, to reduce the communication overhead at the cost of a moderate runtime. Moreover, extensive evaluations conducted on real-world medical datasets demonstrate that our scheme realizes up to 27% accuracy gains.
Rui Zhang 0086, Hongwei Li 0001, Meng Hao 0001, Hanxiao Chen 0001, Yuan Zhang 0006
ICC3
2022 Iron: Private Inference on Transformers
abstract
We initiate the study of private inference on Transformer-based models in the client-server setting, where clients have private inputs and servers hold proprietary models. Our main contribution is to provide several new secure protocols for matrix multiplication and complex non-linear functions like Softmax, GELU activations, and LayerNorm, which are critical components of Transformers. Specifically, we first propose a customized homomorphic encryption-based protocol for matrix multiplication that crucially relies on a novel compact packing technique. This design achieves $\sqrt{m} \times$ less communication ($m$ is the number of rows of the output matrix) over the most efficient work. Second, we design efficient protocols for three non-linear functions via integrating advanced underlying protocols and specialized optimizations. Compared to the state-of-the-art protocols, our recipes reduce about half of the communication and computation overhead. Furthermore, all protocols are numerically precise, which preserve the model accuracy of plaintext. These techniques together allow us to implement \Name, an efficient Transformer-based private inference framework. Experiments conducted on several real-world datasets and models demonstrate that \Name achieves $3 \sim 14\times$ less communication and $3 \sim 11\times$ less runtime compared to the prior art.
Meng Hao 0001, Hongwei Li 0001, Hanxiao Chen 0001, Pengzhi Xing, Guowen Xu, Tianwei Zhang 0004
NeurIPS1
2022 Practical Membership Inference Attack Against Collaborative Inference in Industrial IoT
abstract
The effectiveness of state-of-the-art deep learning (DL) models has empowered the development of industrial Internet of things (IIoT). Recently, considering resource-constrained and privacy-required IIoT devices, collaborative inference has been proposed, which splits DL models and deploys them in IIoT devices and an edge server separately. However, in this article, we argue that there are still severe privacy vulnerabilities in collaborative inference systems. And we devise the first membership inference attack (MIA) against collaborative inference, to infer whether a particular data sample is used for training the model of IIoT systems. Existing MIAs either assume full access to the systems’ APIs or availability of the target model's parameters, which is not applicable in realistic IIoT environments. In contrast to prior works, we proposetransfer-inheritshadow learning and thus relax these key assumptions. We evaluate our attack on different datasets and various settings, and the results show it has high effectiveness.
Hanxiao Chen 0001, Hongwei Li 0001, Guishan Dong, Meng Hao 0001, Guowen Xu, Zhe Liu 0001
IEEE Trans. Ind. Informatics4
2021 Efficient, Private and Robust Federated Learning
abstract
Federated learning (FL) has demonstrated tremendous success in various mission-critical large-scale scenarios. However, such promising distributed learning paradigm is still vulnerable to privacy inference and byzantine attacks. The former aims to infer the privacy of target participants involved in training, while the latter focuses on destroying the integrity of the constructed model. To mitigate the above two issues, a few works recently explored unified solutions by utilizing generic secure computation techniques and common byzantine-robust aggregation rules, but there are two major limitations: 1) they suffer from impracticality due to efficiency bottlenecks, and 2) they are still vulnerable to various types of attacks because of model incomprehensiveness.
Meng Hao 0001, Hongwei Li 0001, Guowen Xu, Hanxiao Chen 0001, Tianwei Zhang 0004
ACSAC1
2021 Towards Lightweight and Efficient Distributed Intrusion Detection Framework
abstract
Federated learning (FL), as a promising distributed learning paradigm, has put many efforts into distributed intrusion detection systems (IDS), for defending against various malicious attacks, such as SQL injection and DDoS attacks. Compared with traditional IDS based on centralized deep learning (DL), FL-based solutions require not to share users' raw data while yielding better detection performance. However, state-of-the-art FL-based methods still suffer from two key limitations: 1) insufficient detection performance on non-independent and identically distributed (non-IID) data, and 2) high communication and computational overheads due to the utilization of large-scale neural network models. In this paper, we propose a lightweight collaborative intrusion detection framework, called CoLGBM, the first of its kind in the regime of decentralized IDS, where decision tree and light gradient boosting machine (LGBM) are combined for constructing the detection scheme. The main insight is that through combining user-trained decision trees (each user's decision tree is derived from its own data with unique distribution), our framework can perform effectively on non-IID data while working efficiently for handling enormous samples. Compared with the current FL-based methods, our CoLGBM achieves higher accuracy and lower overhead on both IID and non-IID data. Extensive experiment results demonstrate our scheme with high-level performance.
Shuai Yuan 0009, Hongwei Li 0001, Rui Zhang 0086, Meng Hao 0001, Rongxing Lu
GLOBECOM4
2021 Enhanced Mixup Training: a Defense Method Against Membership Inference Attack
Zongqi Chen, Hongwei Li 0001, Meng Hao 0001, Guowen Xu
ISPEC3
2020 Privacy-aware and Resource-saving Collaborative Learning for Healthcare in Cloud Computing
abstract
Electronic health records (EHR), generated in healthcare, contain extensive digital information, such as diagnoses, medications and complications. Recently, many studies have focused on constructing deep learning (DL) models with EHR data to improve the quality of healthcare services. However, in traditional centralized training, the collection of EHR causes serious privacy issues due to vulnerable transmission channels and untrusted DL service providers. An alternative that can mitigate the above privacy threat is federated learning (FL). It enables multiple healthcare institutions to learn a global predictive model by exchanging locally calculated updates without disclosing the private dataset. Unfortunately, the latest studies have shown that the local updates still expose sensitive information about the original training data. While several privacy-preserving FL protocols have been proposed, few prior works focused on energy consumption issues. Specifically, local training requires extensive computational resources, which is prohibitively expensive for resource-limited institutions. To overcome the above problems, we propose PRCL, a Privacy-aware and Resource-saving Collaborative Learning protocol. To reduce the local computational overhead, we design a novel model splitting method that partitions the neural network into three parts and outsources the computationally large middle part to cloud servers. By using the lightweight data perturbation and packed partially homomorphic encryption, PRCL protects the privacy of the original data and labels, as well as the parameters of the model. Moreover, we analyze the security of the proposed protocol, and demonstrate the superior performance of PRCL in terms of accuracy and efficiency.
Meng Hao 0001, Hongwei Li 0001, Guowen Xu, Zhe Liu 0001, Zongqi Chen
ICC1
2020 Efficient and Privacy-Enhanced Federated Learning for Industrial Artificial Intelligence
abstract
By leveraging deep learning-based technologies, industrial artificial intelligence (IAI) has been applied to solve various industrial challenging problems in Industry 4.0. However, for privacy reasons, traditional centralized training may be unsuitable for sensitive data-driven industrial scenarios, such as healthcare and autopilot. Recently, federated learning has received widespread attention, since it enables participants to collaboratively learn a shared model without revealing their local data. However, studies have shown that, by exploiting the shared parameters adversaries can still compromise industrial applications such as auto-driving navigation systems, medical data in wearable devices, and industrial robots' decision making. In this article, to solve this problem, we propose an efficient and privacy-enhanced federated learning (PEFL) scheme for IAI. Compared with existing solutions, PEFL is noninteractive, and can prevent private data from being leaked even if multiple entities collude with each other. Moreover, extensive experiments with real-world data demonstrate the superiority of PEFL in terms of accuracy and efficiency.
Meng Hao 0001, Hongwei Li 0001, Xizhao Luo, Guowen Xu, Haomiao Yang, Sen Liu 0007
IEEE Trans. Ind. Informatics1
2019 Towards Efficient and Privacy-Preserving Federated Deep Learning
abstract
Deep learning has been applied in many areas, such as computer vision, natural language processing and emotion analysis. Differing from the traditional deep learning that collects users' data centrally, federated deep learning requires participants to train the networks on private datasets and share the training results, and hence has more gratifying efficiency and stronger security. However, it still presents some privacy issues since adversaries can deduce users' privacy from local outputs, such as gradients. While the problem of private federated deep learning has been an active research issue, the latest research findings are still inadequate in terms of security, accuracy and efficiency. In this paper, we propose an efficient and privacy-preserving federated deep learning protocol based on stochastic gradient descent method by integrating the additively homomorphic encryption with differential privacy. Specifically, users add noises to each local gradients before encrypting them to obtain the optical performance and security. Moreover, our scheme is secure to honest-but-curious server setting even if the cloud server colludes with multiple users. Besides, our scheme supports federated learning for large-scale users scenarios and extensive experiments demonstrate our scheme has high efficiency and high accuracy compared with non-private model.
Meng Hao 0001, Hongwei Li 0001, Guowen Xu, Sen Liu 0007, Haomiao Yang
ICC1