Yang Liu 0118

dblp:51/3710-118 · DBLP profile ↗
← Back
46ranked-venue papers
9as first author
36since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 18 · 5 first-author · 16 since 2021Computer networks · 9 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Systems, architecture and hardware · 3 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Improving Sustainability of Adversarial Examples in Class-Incremental Learning
abstract
Current adversarial examples (AEs) are typically designed for static models. However, with the wide application of Class-Incremental Learning (CIL), models are no longer static and need to be updated with new data distributed and labeled differently from the old ones. As a result, existing AEs often fail after CIL updates due to significant domain drift. In this paper, we propose SAE to enhance the sustainability of AEs against CIL. The core idea of SAE is to enhance the robustness of AE semantics against domain drift by making them more similar to the target class while distinguishing them from all other classes. Achieving this is challenging, as relying solely on the initial CIL model to optimize AE semantics often leads to overfitting. To resolve the problem, we propose a Semantic Correction Module. This module encourages the AE semantics to be generalized, based on a generative model capable of producing universal semantics. Additionally, it incorporates the CIL model to correct the optimization direction of the AE semantics, guiding them closer to the target class. To further reduce fluctuations in AE semantics, we propose a Filtering-and-Augmentation Module, which first identifies non-target examples with target-class semantics in the latent space and then augments them to foster more stable semantics. Comprehensive experiments demonstrate that SAE outperforms baselines by an average of 31.28% when updated with a 9-fold increase in the number of classes.
Taifeng Liu, Xinjing Liu, Liangqiu Dong, Yang Liu 0118, Yilong Yang 0004, Zhuo Ma 0001
AAAI4
2026 AttMark: Attention Based Model Watermarking Against Stealing Attacks
abstract
Model watermarking is a technique that embeds identification information as watermarks to verify model ownership and protect model priority against model stealing (MS) attacks. Watermark is a type of external knowledge which typically make a model sensitive to a specific trigger pattern, causing it to misclassify patterns to a targeted class. However, current fixed form of trigger pattern makes watermarks easy to be recovered by adversaries, thus compromising their secrecy. In this paper, we propose a new approach, named AttMark, which can generate unique patterns for each input via a group of generators. The application of generators adds randomness to trigger patterns by embedding characters into samples in various ways. Therefore, it challenges the convergence of watermark recovering algorithms of adversaries. Nevertheless, random trigger patterns render them more difficult to be embedded, making it even more challenging to transfer watermarks to stolen models. Thus, we design attention-based watermarks that leverage the characteristic of attention transferring in MS attacks. By minimizing the attention deviation caused by random trigger patterns, we enable the stolen model to learn watermarks simultaneously with the primary task. AttMark is evaluated on three major MS attacks and the watermark validation rate is tested against recovering and removal attacks. The results show that our watermark cannot be recovered by adversaries and has a$30\%$stronger transferability compared to prior works. Our code will be available11https://github.com/LiuJingjinga/AttMark.git.
Xinjing Liu, Zhuo Ma 0001, Yang Liu 0118, Taifeng Liu, Zhan Qin
IEEE Trans. Dependable Secur. Comput.3
2026 Urey-ML: A Machine Learning-Based Distance Deception Attack Against Apple UWB Interaction Frameworks
abstract
Ultra-Wideband (UWB) technology has recently emerged as a transformative enabler of high-precision positioning systems. Despite its growing adoption across diverse applications, prior studies have claimed several successful distance deception attacks against UWB. To address heightened security concerns, companies like Apple introduce the ranging-awareness defense mechanism into their new version of the UWB interaction frameworks, which is proven to be effective against most known attacks. In this paper, we critically focus on the design flaws of state-of-the-art UWB interaction frameworks and propose Urey-ML, a novel machine learning-based UWB distance deception attack targeting UWB systems. To the best of our knowledge, this is the first attack capable of circumventing the defense mechanisms implemented in Apple’s UWB Nearby Interaction Framework (ANIF). Specifically, Urey-ML is built upon two critical breakthroughs. First, through network packet analysis, we discover that ANIF leaves a crucial message for key negotiation in an unprotected state. This vulnerability enables Urey-ML to bypass the encryption protection implemented by standard UWB systems. Second, to break the ranging-awareness defense, Urey-ML involves a reinforcement learning-based algorithm to optimize attack parameters. By leveraging this approach, Urey-ML can automatically and craftily generate attack signals that mimic the variations typically caused by normal human movement. Our experiments on commercial-off-the-shelf UWB products show that Urey-ML achieves centimeter-level UWB distance deception, with more than 25.79% signals circumventing the defense check of the victim device, which is only 0.56% (or failed) in prior works.
Yang Liu 0118, Man Sun, Xinjing Liu, Yong Zeng 0002, Jiayu Jin, Zhuo Ma 0001
IEEE Trans. Inf. Forensics Secur.1
2026 PROTheft: A Projector-Based Model Extraction Attack in the Physical World
Xinjing Liu, Yilong Yang 0004, Taifeng Liu, Leo Yu Zhang, Yanjun Zhang 0002, Yang Liu 0118, Zhuo Ma 0001
IEEE Trans. Inf. Forensics Secur.7
2026 chamaeleon: Backdoor Attacks Against Vertical Federated Learning for Tabular Data
abstract
Vertical federated learning (VFL) has made significant strides in enhancing data privacy and security for cross-silo applications. However, despite its benefits, VFL remains vulnerable to emerging security threats, particularly backdoor attacks. While most existing research on VFL backdoor attacks has focused on image and natural language processing tasks, the security of tabular data—commonly used in high-risk domains such as finance and healthcare—has been largely overlooked. In this paper, we introduce chamaeleon, a novel backdoor attack targeting VFL for tabular data. Our approach achieves two key advancements. First, to address the challenge of restricted label access in VFL, chamaeleon employs a two-step inference method to extract label information. This method combines a label classifier with a top-kconfidence filtering mechanism, enabling the precise identification of target-label samples (i.e., backdoored samples) with a precision of approximately 99.85%. Second, to overcome the limitations of fixed trigger patterns, which can disrupt the semantic integrity of tabular data (e.g., altering “male” to “pregnant”), chamaeleon introduces a dynamic trigger design. Each backdoored sample is injected with a unique trigger, generated by a transformer-based model inspired by large language models, ensuring semantic consistency. Additionally, a one-on-two adversarial game is implemented to optimize the generator’s performance with limited training data. Extensive evaluations across six models and six datasets demonstrate the effectiveness of our proposed attack. We also examine various factors that could influence the attack success and systematically analyze potential defense mechanisms to mitigate this newly identified threat.
Yilong Yang 0004, Yong Zeng 0002, Shangze Li, Yang Liu 0118, Zhuo Ma 0001
IEEE Trans. Inf. Forensics Secur.6
2025 Harnessing Large Models, Distilling to Small: Localized Deployment for Accurate Medical Prescription Diagnostic Inference
abstract
Diagnostic errors impose substantial healthcare costs. To address this, we propose BELL, a framework that leverages LLMs for data augmentation and distills knowledge into compact BERT models for efficient deployment. Our twostage framework first standardizes non-uniform clinical terms using fine-tuned BERT models, followed by multi-label disease prediction incorporating prescription data. Experiments on realworld anonymized data demonstrate BELL achieves 94.27 % standardization accuracy and improves diagnostic F1-score from 0.45 to 0.73, with 0.678 s average inference time.
Runjie Zhou, Pujun Feng, Yang Liu 0118, Yuxue Qi, Tong Yang 0003, Bin Cui 0001
BIBM4
2025 SafeLead: Detecting and Excluding Random STS Attack in UWB Ranging System
Zhuo Ma 0001, Jiayu Jin, Yang Liu 0118, Yilong Yang 0004, Xinjing Liu, Teng Li 0003, Junwei Zhang 0001, Jianfeng Ma 0001
INFOCOM3
2025 L-HAWK: A Controllable Physical Adversarial Patch Against a Long-Distance Target
Taifeng Liu, Yang Liu 0118, Zhuo Ma 0001, Tong Yang 0003, Xinjing Liu, Teng Li 0003, Jianfeng Ma 0001
NDSS2
2025 Efficient 2PC for Constant Round Secure Equality Testing and Comparison
Tianpei Lu, Bingsheng Zhang, Zhuo Ma 0001, Yang Liu 0118, Kui Ren 0001, Chun Chen 0001
USENIX Security Symposium6
2025 On the Trustworthiness Landscape of State-of-the-art Generative Models: A Survey and Outlook
Mingyuan Fan 0003, Chengyu Wang 0001, Cen Chen 0001, Yang Liu 0118, Jun Huang 0007
Int. J. Comput. Vis.4
2025 S-Teapot: Swift and Efficient Defense Against Patch-Based Backdoor Attack
abstract
Recent studies emphasize the serious threat posed by backdoor attacks when training deep models on data from untrustworthy sources. Despite the emergence of various backdoor attack paradigms, the patch-based approach stands out as the most sought-after and effective method of poisoning. However, current defenses against such attacks often exhibit rudimentary and highly inefficient, sometimes necessitating days for implementation. To mitigate this, we propose a swift and efficient defense against patch-based backdoor attacks, calledS-teapot.S-teapotrapidly identifies whether an untrusted dataset has been backdoored and determines the backdoored labels based on the model's high confidence in the poisoning sample and the consistency of the backdoor pixels.S-teapotoutperforms existing backdoor attack detection schemes by a speedup factor ranging from 30 to 259. Furthermore, we leverage the abnormality of the backdoor pixels to reverse the backdoor trigger, resulting in a similarity increase of 0.6 to 32 times compared to existing methods. To obtain a clean model,S-teapotaccurately localizes poisoning samples through similarity calculations, with nearly 100% precision. Leveraging the precision of the reverse triggers,S-teapotemploys an inpaint method to convert the poisoning samples into clean ones, yielding up to 8.16% improvement in accuracy.
Yilong Yang 0004, Zhuo Ma 0001, Yihua Li, Yang Liu 0118, Xinjing Liu, Jianfeng Ma 0001
IEEE Trans. Dependable Secur. Comput.4
2025 RAPOO: An Efficient Privacy-Preserving Facial Expression Recognition via Mobile Crowdsensing
abstract
Facial expression recognition is a technology that involves analyzing and interpreting human facial expressions to determine individual expressions or states. Mobile crowdsensing (MCS), a promising sensing paradigm, makes it easy to capture facial images and benefits facial expression recognition. Existing inference models for facial expression recognition usually rely on facial feature vectors or facial images, increasing privacy concerns about expression. For this reason, this paper proposes a privacy-preserving facial expression recognition scheme through MCS, named RAPOO, which falls in a client-server architecture. Roughly speaking, a user captures facial images using mobile devices and requests a recognition service provided by a cloud computing center. To protect the privacy of expressions, our approach focuses on designing secure computation protocols required by facial expression recognition necessarily, such as secure vector distance calculation and secure top-$k$query. These protocols enable facial expression recognition over encrypted data directly. To speed up the recognition and store encrypted feature vectors, a$k$-D tree data structure is introduced. The security analysis confirms that RAPOO effectively preserves the confidentiality of personal expressions. Extensive experimental evaluations show that our solution obtains a three-order-of-magnitude speedup in terms of computational overhead compared with the state-of-the-art.
Bowen Zhao 0001, Yang Xiao 0014, Yang Liu 0118, Qingqi Pei, Yulong Shen 0001
IEEE Trans. Mob. Comput.4
2024 Updates Leakage Attack Against Private Graph Split Learning
Zhuo Ma 0001, Yang Liu 0118, Xinjing Liu, Beiwei Yang, Jianfeng Ma 0001
ICA3PP (2)3
2024 Need for Speed: Taming Backdoor Attacks with Speed and Precision
abstract
Modern deep neural network models (DNNs) require extensive data for optimal performance, prompting reliance on multiple entities for the acquisition of training datasets. One prominent security threat is backdoor attacks where the adversary party poisons a small subset of training datasets to implant a backdoor into the model, leading to misclassifications during runtime for triggered samples. To mitigate the attack, many defense methods have been proposed, such as detecting and removing poisoned samples or rectifying trojaned model weights in victim DNNs. However, existing approaches suffer from notable inefficiency as they are faced with large-scale training datasets, consequently rendering these defenses impractical in the real world. In this paper, we propose a lightweight backdoor identification and removal scheme, called ReBack. In this scheme, ReBack first extracts a subset of suspicious and benign samples, and then, proceeds with a "averaging and differencing" based method to identify target label(s). Next, leveraging the identification results, ReBack invokes a novel reverse engineering method to recover the exact trigger using only basic arithmetic atoms. Our experiments demonstrate that, for ImageNet with 750 labels, ReBack can defend against backdoor attacks in around 2 hours, showcasing a speed improvement of 18.5× to 214× compared to existing methods. For backdoor removal, the attack success rate can be decreased to 0.05% owing to 99% cosine similarity of the reversed triggers. The code is online available.
Zhuo Ma 0001, Yilong Yang 0004, Yang Liu 0118, Tong Yang 0003, Xinjing Liu, Teng Li 0003, Zhan Qin
SP3
2024 Guardian: Guarding against Gradient Leakage with Provable Defense for Federated Learning
abstract
Federated learning is a privacy-focused learning paradigm, which trains a global model with gradients uploaded from multiple participants, circumventing explicit exposure of private data. However, previous research of gradient leakage attacks suggests that gradients alone are sufficient to reconstruct private data, rendering the privacy protection mechanism of federated learning unreliable. Existing defenses commonly craft transformed gradients based on ground-truth gradients to obfuscate the attacks, but often are less capable of maintaining good model performance together with satisfactory privacy protection. In this paper, we propose a novel yet effective defense framework named guarding against gradient leakage (Guardian) that produces transformed gradients by jointly optimizing two theoretically-derived metrics associated with gradients for performance maintenance and privacy protection. In this way, the transformed gradients produced via Guardian can achieve minimal privacy leakage in theory with the given performance maintenance level. Moreover, we design an ingenious initialization strategy for faster generation of transformed gradients to enhance the practicality of Guardian in real-world applications, while demonstrating theoretical convergence of Guardian to the performance of the global model. Extensive experiments on various tasks show that, without sacrificing much accuracy, Guardian can effectively defend state-of-the-art gradient leakage attacks, compared with the slight effects of baseline defense approaches.
Mingyuan Fan 0003, Yang Liu 0118, Cen Chen 0001, Chengyu Wang 0001, Minghui Qiu, Wenmeng Zhou
WSDM2
2024 Mitigate noisy data for smart IoT via GAN based machine unlearning
Zhuo Ma 0001, Yilong Yang 0004, Yang Liu 0118, Xinjing Liu, Jianfeng Ma 0001
Sci. China Inf. Sci.3
2024 Efficient and self-recoverable privacy-preserving k-NN classification system with robustness to network delay
Jinhai Zhang, Junwei Zhang 0001, Zhuo Ma 0001, Yang Liu 0118, XinDi Ma, Jianfeng Ma 0001
J. Syst. Archit.4
2024 Effectively Improving Data Diversity of Substitute Training for Data-Free Black-Box Attack
abstract
Recent substitute training methods have utilized the concept of Generative Adversarial Networks (GANs) to implement data-free black-box attacks. Specifically, in designing the generators, the substitute training methods use a similar structure to the generators in GANs. However, this design approach ignores the potential situation that the generators in GANs operate under real data supervision, while the generators in substitute training methods lack such supervision. This difference in data-supervised conditions constrain the diversity of data generated by the substitute training methods, resulting in inadequate data to support effective training of the substitute model. This impacts the substitute model's ability to attack the target model further. Consequently, to solve the above issues, we propose three strategies to improve the attack success rates. For the generator, we first propose a dense projection space that projects the input noise into various latent feature spaces to diversify feature information. Then, we introduce a novel disguised natural color mode. This mode improves information exchange between the generator's output layer and previous layers, allowing for more diverse generated data. Besides, we present a regularization method for the substitute model, called noise-based balanced learning, to prevent the potential risk of overfitting due to the lack of diversity of the generated data. In the experimental analysis, extensive experiments are conducted to validate the effectiveness of these proposed strategies.
Yang Wei 0002, Zhuo Ma 0001, Zhuoran Ma 0002, Zhan Qin, Yang Liu 0118, Bin Xiao 0002, Xiuli Bi, Jianfeng Ma 0001
IEEE Trans. Dependable Secur. Comput.5
2024 FlGan: GAN-Based Unbiased Federated Learning Under Non-IID Settings
abstract
Federated Learning (FL) suffers from low convergence and significant accuracy loss due to local biases caused by non-Independent and Identically Distributed (non-IID) data. To enhance the non-IID FL performance, a straightforward idea is to leverage the Generative Adversarial Network (GAN) to mitigate local biases using synthesized samples. Unfortunately, existing GAN-based solutions have inherent limitations, which do not support non-IID data and even compromise user privacy. To tackle the above issues, we propose a GAN-based unbiased FL scheme, calledFlGan, to mitigate local biases using synthesized samples generated by GAN while preserving user-level privacy in the FL setting. Specifically,FlGanfirst presents a federated GAN algorithm using the divide-and-conquer strategy that eliminates the problem of model collapse in non-IID settings. To guarantee user-level privacy,FlGanthen exploits Fully Homomorphic Encryption (FHE) to design the privacy-preserving GAN augmentation method for the unbiased FL. Extensive experiments show thatFlGanachieves unbiased FL with$10\%-60\%$accuracy improvement compared with two state-of-the-art FL baselines (i.e., FedAvg and FedSGD) trained under different non-IID settings. The FHE-based privacy guarantees only cost about 0.53% of the total overhead inFlGan.
Zhuoran Ma 0002, Yang Liu 0118, Yinbin Miao, Guowen Xu, Ximeng Liu, Jianfeng Ma 0001, Robert H. Deng
IEEE Trans. Knowl. Data Eng.2
2024 Multi-Party Private Edge Computing for Collaborative Quantitative Exposure Detection of Endemic Diseases
abstract
Facing the global threat of endemic diseases, utilizing edge computing for exposure detection enables efficient monitoring of the dynamic distribution of infected patient groups across regions, enhancing the management and control of these diseases. Employing the quantitative exposure detection of endemic diseases, regions seek to reconcile patient information collected through mobile devices, aiming to obtain statistical and analytical results based on the intersection of patient lists. In this paper, we propose a privacy-preserving scheme for the collaborative quantitative exposure detection of endemic diseases, which ensures each region to only learn the statistical results, without any information about other regions' datasets. Our scheme is fundamentally achieved through Circuit-based Private Set Intersection (Circuit-PSI) that can compute functions over the set intersection without disclosing the intersection itself. However, the state-of-the-art solution involves a laborious process in which one party iteratively compares its elements with those of others, which leads to a significantly high communication complexity. Therefore, we introduce a novel multi-party protocol that can diminish the communication overhead of circuit-PSI through a skillful decoupling of the comparison complexity from the number of parties. To achieve this, we design a multiparty oblivious encoding scheme, which can prevent any party from inferring any private info through the encoded data. By filtering out the repeated elements, the comparison complexity is independent of the number of parties. Furthermore, to address scenarios involving patient information with additional attributes, we extend our protocol to include payloads by developing a lightweight multiparty data mapping algorithm. Our extensive experiments show that compared to prior works, our protocol achieves a substantial reduction in communication overhead by 6.4×, and runs 1.2× faster in the LAN setting and 3.1× in the WAN setting.
Zhuo Ma 0001, Yang Liu 0118, Teng Li 0003, Zuobin Ying, Bingsheng Zhang
IEEE Trans. Mob. Comput.3
2024 A Privacy-Preserving Computation Framework for Multisource Label Propagation Services
abstract
Multisource Private Label Propagation (MPLP) is designed for different organizations to collaboratively predict labels of unlabeled nodes through iterative propagation and label updates without revealing sensitive information. Aside from the privacy of the origin data, in some statistical prediction services, it is only needed to learn about the statistical results and concrete prediction results for the abnormal nodes. To do it, we first design a basic MPLP scheme,PriLP, to meet the requirements of the privacy of origin data and the concrete prediction of normal nodes. However, our basic achievement ofPriLPrelies heavily on Additive Homomorphic Encryption (AHE) due to the sparse graph representation in label propagation. To diminish reliance on AHE, our optimization facilitates data encryption in a more compact representation, resulting in encryption times that scale linearly with the number of graph nodes. Our experiments showPriLPclosely matches plain-label propagation within$\leq 0.7\%$difference in accuracy, and the optimizations lead to up to$22.63\times$faster execution and$1.83\times$less communication than the basic implement.
Tanren Liu, Zhuo Ma 0001, Yang Liu 0118, Bingsheng Zhang, Jianfeng Ma 0001
IEEE Trans. Serv. Comput.3
2023 Secondary Labeling: A Novel Labeling Strategy for Image Manipulation Detection
abstract
Image manipulation detection methods typically rely on a binary annotation called Primary Labeling (PrLa) to identify tampered and authentic regions in a tampered image. However, PrLa only focuses on the difference between authentic and tampered regions, ignoring the distinctions among tampered regions in different images. This transforms the task of image manipulation detection into salient object detection, with the goal shifting towards identifying the most attention-grabbing objects in images. To address this issue, this paper proposes a novel labeling strategy called Secondary Labeling (SeLa). SeLa generates a query table containing multiple tampered categories and randomly reassigns these tampered classes to different types of tampered data, effectively improving the detection performance of models by refocusing the differences among the various data. Additionally, to further improve the detection performance, this paper introduces an Adaptive Label Smoothing (ALS) regularization method. This method addresses the loss of correlation among tampered classes in SeLa caused by the one-hot encoding method. Experimental results show that compared with PrLa, SeLa not only improves the performance of detection models by up to 17%, but also enhances the robustness and convergence rate.
Yang Wei 0002, Bin Xiao 0002, Xiuli Bi, Zhuoran Ma 0002, Yang Liu 0118, Zhuo Ma 0001
ACM Multimedia5
2023 Sniffer: A Novel Model Type Detection System against Machine-Learning-as-a-Service Platforms
abstract
Recent works explore several attacks against Machine-Learning-as-a-Service (MLaaS) platforms (e.g., the model stealing attack), allegedly posing potential real-world threats beyond viability in laboratories. However, hampered by model-type-sensitive , most of the attacks can hardly break mainstream real-world MLaaS platforms. That is, many MLaaS attacks are designed against only one certain type of model, such as tree models or neural networks. As the black-box MLaaS interface hides model type info, the attacker cannot choose a proper attack method with confidence, limiting the attack performance. In this paper, we demonstrate a system, named Sniffer, that is capable of making model-type-sensitive attacks "great again" in real-world applications. Specifically, Sniffer consists of four components: Generator, Querier, Probe, and Arsenal. The first two components work for preparing attack samples. Probe, as the most characteristic component in Sniffer, implements a series of self-designed algorithms to determine the type of models hidden behind the black-box MLaaS interfaces. With model type info unraveled, an optimum method can be selected from Arsenal (containing multiple attack methods) to accomplish its attack. Our demonstration shows how the audience can interact with Sniffer in a web-based interface against five mainstream MLaaS platforms.
Zhuo Ma 0001, Yilong Yang 0004, Bin Xiao 0002, Yang Liu 0118, Xinjing Liu, Zhuoran Ma 0002, Tong Yang 0003
Proc. VLDB Endow.4
2023 Outsourced Privacy-Preserving Data Alignment on Vertically Partitioned Database
abstract
In the context of real-world secure outsourced computations, private data alignment has been always the essential preprocessing step. However, current private data alignment schemes, mainly circuit-based, suffer from high communication overhead and often need to transfer potentially gigabytes of data. In this paper, we propose a lightweight private data alignment protocol (called SC-PSI) that can overcome the bottleneck of communication. Specifically, SC-PSI involves four phases of computations, including data preprocessing, data outsourcing, private set member (PSM) evaluation and circuit computation (CC). Like prior works, the major overhead of SC-PSI mainly lies in the latter two phases. The improvement is SC-PSI utilizes the function secret sharing technique to develop the PSM protocol, which avoids the multiple rounds of communication to compute intersection set members. Moreover, benefited from our specially designed PSM protocol, SC-PSI does not to execute complex secure comparison circuits in the CC phase. Experimentally, we validate that compared to prior works, SC-PSI can save around 61.39% running time and 89.61% communication overhead.
Cui Hu, Bin Xiao 0002, Yang Liu 0118, Teng Li 0003, Zhuo Ma 0001, Jianfeng Ma 0001
IEEE Trans. Big Data4
2023 Learn to Forget: Machine Unlearning via Neuron Masking
abstract
Nowadays, machine learning models, especially neural networks, have became prevalent in many real-world applications. These models are trained based on a one-way trip from user data: as long as users contribute their data, there is no way to withdraw. To this end,machine unlearningbecomes a popular research topic, which allows the model trainer to unlearn unexpected data from a trained machine learning model. In this article, we propose the first uniform metric called forgetting rate to measure the effectiveness of a machine unlearning method. It is based on the concept of membership inference and describes the transformation rate of the eliminated data from “memorized” to “unknown” after conducting unlearning. We also propose a novel unlearning method calledForsaken. It is superior to previous work in either utility or efficiency (when achieving the same forgetting rate). We benchmarkForsakenwith eight standard datasets to evaluate its performance. The experimental results show that it can achieve more than 90% forgetting rate on average and only causeless than 5% accuracy loss.
Zhuo Ma 0001, Yang Liu 0118, Ximeng Liu, Jian Liu 0012, Jianfeng Ma 0001, Kui Ren 0001
IEEE Trans. Dependable Secur. Comput.2
2023 DivTheft: An Ensemble Model Stealing Attack by Divide-and-Conquer
abstract
Recently, model stealing attacks are widely studied but most of them are focused on stealing a single non-discrete model, e.g., neural networks. For ensemble models, these attacks are either non-executable or suffer from intolerant performance degradation due to the complex model structure (multiple sub-models) and the discreteness possessed by the sub-model (e.g., decision trees). To overcome the bottleneck, this paper proposes a divide-and-conquer strategy called DivTheft to formulate the model stealing attack to common ensemble models by combining active learning (AL). Specifically, based on the boosting learning concept, we divide a hard ensemble model stealing task into multiple simpler ones about single sub-model stealing. Then, we adopt AL to conquer the data-free sub-model stealing task. During the process, the current AL algorithm easily causes the stolen model to be biased because of ignoring the past useful memories. Thus, DivTheft involves a newly designed uncertainty sampling scheme to filter reusable samples from the previously used ones. Experiments show that compared with the prior work, DivTheft can save almost 50% queries while ensuring a competitive agreement rate to the victim model.
Zhuo Ma 0001, Xinjing Liu, Yang Liu 0118, Ximeng Liu, Zhan Qin, Kui Ren 0001
IEEE Trans. Dependable Secur. Comput.3
2023 Forward/Backward and Content Private DSSE for Spatial Keyword Queries
abstract
Spatial keyword queries are attractive techniques that have been widely deployed in real-life applications in recent years, such as social networks and location-based services. However, existing solutions neither support dynamic update nor satisfy the privacy requirements in real applications. In this article, we investigate the problem of Dynamic Searchable Symmetric Encryption (DSSE) for spatial keyword queries. First, we formulate the definition of DSSE for spatial keyword queries (namely, DSSESKQ) and extend the DSSE leakage functions to capture the leakages in DSSESKQ. Then, we present a practical DSSESKQ construction based on geometric prefix encoding inverted-index and encrypted bitmap. Rigorous security analysis proves that our construction can achieve not only forward/backward privacy but content privacy as well, which can resist the most existing leakage-abuse attacks. Evaluation results using real-world datasets demonstrate the efficiency and feasibility of our construction. Comparative analysis reveals that our construction outperforms state-of-the-art schemes in terms of privacy and performance, e.g., our construction is 175x faster than existing schemes with only 51% server storage cost.
Xiangyu Wang 0010, Jianfeng Ma 0001, Ximeng Liu, Yinbin Miao, Yang Liu 0118, Robert H. Deng
IEEE Trans. Dependable Secur. Comput.5
2023 iPrivJoin: An ID-Private Data Join Framework for Privacy-Preserving Machine Learning
abstract
The world has observed an increasing trend in the development of Privacy-Preserving Machine Learning (PPML) for cross-silo collaborative model training over sensitive data. As the first essential step of cross-silo PPML, it is critical that the parties can align their dataset with privacy assurance, i.e.,private data join. However, the existing private data join methods typically leak the ID information in the dataset intersection, which often raises privacy concerns. In this work, we propose iPrivJoin: a novel framework of ID-private data join for PPML. Compared with naively using circuit-based Private Set Intersection (circuit-PSI) for data join, the proposed framework has two advantages. (i) data volume reduction. iPrivJoin utilizes oblivious shuffle to securely trim off the redundant data that is outside the intersection, while the entire dataset needs to be carried to further process in the circuit-PSI based approach. (ii) efficiency improvement. iPrivJoin introduces a new private encoding technique to avoid the expensive circuit evaluation that is needed in circuit-PSI. As a result, compared with directly using circuit-PSI, PPML with iPrivJoin enjoys approximately 3× of speedup. Moreover, we propose a new oblivious shuffle protocol, which may be of independent interest. It achieves 1.44× of speedup to the state-of-the-art in the real-world WAN network setting.
Yang Liu 0118, Bingsheng Zhang, Zhuo Ma 0001, Zecheng Wu
IEEE Trans. Inf. Forensics Secur.1
2023 Reveal Your Images: Gradient Leakage Attack Against Unbiased Sampling-Based Secure Aggregation
abstract
Recently, some Unbiased Gradient Sampling-based (UGS) methods have been proposed to enhance the security and efficiency of federated learning through crafted unbiased random transformation and sampling, such as MinMax Sampling in SIGMOD ’22. In this paper, we propose a novel attack, GLAUS, to show that UGS is not as secure as claimed in these works and is still vulnerable to the gradient leakage attack (GLA). Specifically, we demonstrate an idea to approximately infer the gradient for GLA in the context of the UGS scenario where the real gradient is not available. Once the gradient is approximately obtained, the security of the UGS frameworks is downgraded to that of the original federated learning. The approximate gradient is refined by the following steps: 1)narrow the gradient searching rangeto the finite set; 2)obtain the magnitudeof each gradient value approximately; 3)revise the gradient signs. Versus the failure of existing attacks, extensive experiments on six datasets show that our attack is effective in reconstructing private datapoints with pixel-wise accuracy on four network sizes and three image resolutions. Finally, we show how to defend against GLAUS while maintaining the high efficiency of UGS and only introducing an additional step to hide the sampled gradient indices.
Yilong Yang 0004, Zhuo Ma 0001, Bin Xiao 0002, Yang Liu 0118, Teng Li 0003, Junwei Zhang 0008
IEEE Trans. Knowl. Data Eng.4
2022 SeInspect: Defending Model Stealing via Heterogeneous Semantic Inspection
Xinjing Liu, Zhuo Ma 0001, Yang Liu 0118, Zhan Qin, Junwei Zhang 0001
ESORICS (1)3
2022 Combating False Sense of Security: Breaking the Defense of Adversarial Training Via Non-Gradient Adversarial Attack
abstract
Adversarial training is believed to be the most robust and effective defense method against adversarial attacks. Gradient-based adversarial attack methods are generally adopted to evaluate the effectiveness of adversarial training. However, in this paper, by diving into the existing adversarial attack literature, we find that adversarial examples generated by these attack methods tend to be less imperceptible, which may lead to an inaccurate estimation for the effectiveness of the adversarial training. The existing adversarial attacks mostly adopt gradient-based optimization methods and such optimization methods have difficulties in searching the most effective adversarial examples (i.e., the global extreme points). On the contrast, in this work, we propose a novel Non-Gradient Attack (NGA) to overcome the above-mentioned problem. Extensive experiments show that NGA significantly outperforms the state-of-the-art adversarial attacks on Attack Success Rate (ASR) by 2% ∼ 7%.
Mingyuan Fan 0003, Yang Liu 0118, Cen Chen 0001, Shengxing Yu, Wenzhong Guo, Ximeng Liu
ICASSP2
2022 Backdoor Defense with Machine Unlearning
abstract
Backdoor injection attack is an emerging threat to the security of neural networks, however, there still exist limited effective defense methods against the attack. In this paper, we propose BAERASER, a novel method that can erase the backdoor injected into the victim model through machine unlearning. Specifically, BAERASER mainly implements backdoor defense in two key steps. First, trigger pattern recovery is conducted to extract the trigger patterns infected by the victim model. Here, the trigger pattern recovery problem is equivalent to the one of extracting an unknown noise distribution from the victim model, which can be easily resolved by the entropy maximization based generative model. Subsequently, BAERASER leverages these recovered trigger patterns to reverse the backdoor injection procedure and induce the victim model to erase the polluted memories through a newly designed gradient ascent based machine unlearning method. Compared with the previous machine unlearning solutions, the proposed approach gets rid of the reliance on the full access to training data for retraining and shows higher effectiveness on backdoor erasing than existing fine-tuning or pruning methods. Moreover, experiments show that BAERASER can averagely lower the attack success rates of three kinds of state-of-the-art backdoor attacks by 99% on four benchmark datasets.
Yang Liu 0118, Mingyuan Fan 0003, Cen Chen 0001, Ximeng Liu, Zhuo Ma 0001, Li Wang 0056, Jianfeng Ma 0001
INFOCOM1
2022 A certificateless authentication scheme with fuzzy batch verification for federated UAV network
abstract
Recently, the explosive development of unmanned aerial vehicles (UAVs) promotes its wide application in various services such as package delivery, traffic monitoring. However, due to the high-speed movement, current UAVs usually adopts the unsecure channel without the complicated authentication mechanism to ensure real-time communication. In this paper, we introduce a certificateless authentication scheme with fuzzy batch verification (CLFBV) to achieve once-for-all verification of parallel messages and ensure the real-time secure communication of UAVs. CLBFV defines the error tolerance property for authenticated communication, which allows a tolerance threshold for the messages that are unable to pass authentication. In addition, our proposed scheme is proved to be secure and existentially unforgeable under the chosen message attack and fuzzy identity attack in the random oracle model. The efficiency analysis shows that CLBFV is more efficient and feasible than other existing batch verification schemes.
Junwei Zhang 0001, Yang Liu 0118, Maobin Lu, Zuobin Ying, Jianfeng Ma 0001
Int. J. Intell. Syst.3
2022 Toward Evaluating the Reliability of Deep-Neural-Network-Based IoT Devices
abstract
Nowadays, the impressive performance of deep neural networks (DNNs) greatly advances the development of Internet of Things (IoT) in diverse scenarios. However, the exceptional vulnerability of DNNs to adversarial attack leads IoT devices to be exposed to potential security issues. Up to now, since adversarial training empirically remains robust against gradient-based adversarial attacks, it is believed to be the most effective defense method. In this article, we find that adversarial examples generated by gradient-based adversarial attacks tend to be less imperceptible induced by the gradient-based optimization methods (adopted in the attacks) being difficult on searching the most effective adversarial examples (i.e., the global extreme points), which may lead to an inaccurate estimation for the effectiveness of the adversarial training. To overcome the inherent defect of gradient-based adversarial attacks, we propose a novel adversarial attack named nongradient attack (NGA), of which search strategy is effective but no longer depends on gradients to enhance the threat of adversarial examples. In detail, NGA first initializes the adversarial examples outside, rather than inside, of decision boundary to make them misclassified by the model and then, under without violation of misclassified condition, adjusts the adversarial examples toward the crafted direction to close the original examples. Extensive experiments show that NGA significantly outperforms the state-of-the-art adversarial attacks on attack success rate (ASR) by 2%–7%. Moreover, we propose a new evaluation metric, i.e., composite criterion (CC) based on both ASR and accuracy, to better measure the effectiveness of adversarial training. In the experiments, CC has shown to be a more comprehensive yet appropriate evaluation metric.
Mingyuan Fan 0003, Yang Liu 0118, Cen Chen 0001, Shengxing Yu, Wenzhong Guo, Li Wang 0056, Ximeng Liu
IEEE Internet Things J.2
2022 RevFRF: Enabling Cross-Domain Random Forest Training With Revocable Federated Learning
abstract
Random forest is one of the most heated machine learning tools in a wide range of industrial scenarios. Recently, federated learning enables efficient distributed machine learning without direct revealing of private participant data. In this article, we present a novel framework of federated random forest (RevFRF), and further emphatically discuss the participant revocation problem of federated learning based on RevFRF. Specifically, RevFRF first introduces a suite of homomorphic encryption based secure protocols to implement federated random forest (RF). The protocols cover the whole lifecycle of an RF model, including construction, prediction and participant revocation. Then, referring to the practical application scenarios of RevFRF, the existing federated learning frameworks ignore a fact that even every participant in federated learning cannot maintain the cooperation with others forever. In company-level cooperation, allowing the remaining companies to use a trained model that contains the memories from an off-lying company potentially leads to a significant conflict of interest. Therefore, we propose the revocable federated learning concept and illustrate how RevFRF implements participant revocation in applications. Through theoretical analysis and experiments, we show that the protocols can efficiently implement federated RF and ensure the memories of a revoked participant in the trained RF to be securely removed.
Yang Liu 0118, Zhuo Ma 0001, Yilong Yang 0004, Ximeng Liu, Jianfeng Ma 0001, Kui Ren 0001
IEEE Trans. Dependable Secur. Comput.1
2022 Privacy-Preserving Object Detection for Medical Images With Faster R-CNN
abstract
In this paper, we propose a lightweight privacy-preserving Faster R-CNN framework (SecRCNN) for object detection in medical images. Faster R-CNN is one of the most outstanding deep learning models for object detection. Using SecRCNN, healthcare centers can efficiently complete privacy-preserving computations of Faster R-CNN via the additive secret sharing technique and edge computing. To implement SecRCNN, we design a series of interactive protocols to perform the three stages of Faster R-CNN, namely feature map extraction, region proposal and regression and classification. To improve the efficiency of SecRCNN, we improve the existing secure computation sub-protocols involved in SecRCNN, including division, exponentiation and logarithm. The newly proposed sub-protocols can dramatically reduce the number of messages exchanged during the iterative approximation process based on the coordinate rotation digital computer algorithm. Moreover, the effectiveness, efficiency and security of SecRCNN are demonstrated through comprehensive theoretical analysis and extensive experiments. The experimental findings show that the communication overhead in computing division, logarithm and exponentiation decreases to 36.19%, 73.82% and 43.37%, respectively.
Yang Liu 0118, Zhuo Ma 0001, Ximeng Liu, Siqi Ma 0001, Kui Ren 0001
IEEE Trans. Inf. Forensics Secur.1
2020 Spectrum Privacy Preserving for Social Networks: A Personalized Differential Privacy Approach
Yang Liu 0118, Yong Zeng 0002, Jianfeng Ma 0001
Inscrypt1
2020 Boosting Privately: Federated Extreme Gradient Boosting for Mobile Crowdsensing
abstract
Recently, Google and other 24 institutions proposed a series of open challenges towards federated learning (FL), which include application expansion and homomorphic encryption (HE). The former aims to expand the applicable machine learning models of FL. The latter focuses on who holds the secret key when applying HE to FL. For the naive HE scheme, the server is set to master the secret key. Such a setting causes a serious problem that if the server does not conduct aggregation before decryption, a chance is left for the server to access the user’s update. Inspired by the two challenges, we propose FEDXGB, a federated extreme gradient boosting (XGBoost) scheme supporting forced aggregation. FEDXGB mainly achieves the following two breakthroughs. First, FEDXGB involves a new HE based secure aggregation scheme for FL. By combining the advantages of secret sharing and homomorphic encryption, the algorithm can solve the second challenge mentioned above, and is robust to the user dropout. Then, FEDXGB extends FL to a new machine learning model by applying the secure aggregation scheme to the classification and regression tree building of XGBoost. Moreover, we conduct a comprehensive theoretical analysis and extensive experiments to evaluate the security, effectiveness, and efficiency of FEDXGB. The results indicate that FEDXGB achieves less than 1% accuracy loss compared with the original XGBoost, and can provide about 23.9% runtime and 33.3% communication reduction for HE based model update aggregation of FL.
Yang Liu 0118, Zhuo Ma 0001, Ximeng Liu, Siqi Ma 0001, Surya Nepal, Robert H. Deng, Kui Ren 0001
ICDCS1
2020 PE-HEALTH: Enabling Fully Encrypted CNN for Health Monitor with Optimized Communication
abstract
Cloud-based Convolutional neural network (CNN) is a powerful tool for the healthcare center to provide health condition monitor service. Although the new service has future prospects in the medical, patient's privacy concerns arise because of the sensitivity of medical data. Prior works to address the concern have the following unresolved problems: 1) focus on data privacy but neglect to protect the privacy of the machine learning model itself; 2) introduce considerable communication costs for the CNN inference, which lowers the service quality of the cloud server. To push forward this area, we propose PE-HEALTH, a privacy-preserving health monitor framework that supports fully-encrypted CNN (both input data and model). In PE-HEALTH, the medical Internet of Things (IoT) sensor serves as the health condition data collector. For protecting patient privacy, the IoT sensor additively shares the collected data and uploads the shared data to the cloud server, which is efficient and suited to the energy-limited IoT sensor. To keep model privacy, PE-HEALTH allows the healthcare center to previously deploy, and then, use an encrypted CNN on the cloud server. During the CNN inference process, PE-HEALTH does not need the cloud servers to exchange any extra messages for operating the convolutional operation, which can greatly reduce the communication cost.
Yang Liu 0118, Yilong Yang 0004, Zhuo Ma 0001, Ximeng Liu, Siqi Ma 0001
IWQoS1
2020 LiPSG: Lightweight Privacy-Preserving Q-Learning-Based Energy Management for the IoT-Enabled Smart Grid
abstract
As the largest Internet-of-Things (IoT) deployment in the world, the smart grid implements extremely reduction in the energy dissipation for the operation of the smart city. However, the electricity data produced by the smart grid contain massive sensitive information, such as dispatching instructions and bills. The data are always revealed to cloud servers in the plaintext format for the$Q$-learning-based energy strategy making, which gives the chance for the adversary to abuse the user data. Therefore, in this article, we propose a lightweight privacy-preserving$Q$-learning framework (LiPSG) for the energy management strategy making of the smart grid. Before being sent to the control center, the electricity data of each power supply region in LiPSG are first split into uniformly random secret shares. During completion of the computation task of$Q$-learning, the data are kept in the random share format all the time to avoid the data privacy disclosure. The computation feature is implemented by the newly proposed additive secret-sharing protocols. The edge computing technology is also deployed to further improve efficiency. Moreover, comprehensive theoretic analysis and experiments are given to prove the security and efficiency of LiPSG. Compared with the existing privacy-preserving schemes of the smart grid, LiPSG first provides a general$Q$-learning-based privacy-preserving power strategy making architecture with high efficiency and low-performance loss.
Yang Liu 0118, Zhuo Ma 0001, Ximeng Liu, Jianfeng Ma 0001
IEEE Internet Things J.2
2020 Privacy-preserving federated k-means for proactive caching in next generation cellular networks
Yang Liu 0118, Zhuo Ma 0001, Zheng Yan 0002, Ximeng Liu, Jianfeng Ma 0001
Inf. Sci.1
2020 A machine learning-based scheme for the security analysis of authentication and key agreement protocols
Zhuo Ma 0001, Yang Liu 0118, Haoran Ge, Meng Zhao 0001
Neural Comput. Appl.2
2020 EmIr-Auth: Eye Movement and Iris-Based Portable Remote Authentication for Smart Grid
abstract
With the development of Industry 4.0, the communication of smart grid has recently been taken seriously to ensure secure communication between operator and control center. However, the authentication process between them faces many challenges. Once the attacker successfully authenticated in the control center, the privacy data in the smart grid may leak and cause irreparable damage to the user. In addition, operator authentication is one of the most basic and crucial processes. Therefore, we propose theeye-movement and iris recognition based authentication (EmIr-Auth), a novel biometrics-based remote operator authentication scheme.EmIr-Authuses the recorded eye-movement trajectory and randomly selected iris image to authenticate operators, which is beneficial in that it is able to get rid of many cryptographic computations, as well as the need to minimize message exchange. Furthermore, except for a high-resolution camera, we do not require any additional biometric sensors in this scheme. Using the Burrows–Abadi–Needham logic, in this article, we demonstrate that our scheme provides secure authentication. Moreover, we analyze the attacks thatEmIr-Authcan resist by informal security analysis. Experimental results show thatEmIr-Authis efficient enough to deploy on portable devices and reduce the overhead of authentication procedure.
Zhuo Ma 0001, Yilong Yang 0004, Ximeng Liu, Yang Liu 0118, Siqi Ma 0001, Kui Ren 0001
IEEE Trans. Ind. Informatics4
2019 An empirical study of SMS one-time password authentication in Android apps
abstract
A great quantity of user passwords nowadays has been leaked through security breaches of user accounts. To enhance the security of the Password Authentication Protocol (PAP) in such circumstance, Android app developers often implement a complementary One-Time Password (OTP) authentication by utilizing the short message service (SMS). Unfortunately, SMS is not specially designed as a secure service and thus an SMS One-Time Password is vulnerable to many attacks. To check whether a wide variety of currently used SMS OTP authentication protocols in Android apps are properly implemented, this paper presents an empirical study against them. We first derive a set of rules from RFC documents as the guide to implement secure SMS OTP authentication protocol. Then we implement an automated analysis system, AUTH-EYE, to check whether a real-world OTP authentication scheme violates any of these rules. Without accessing server source code, AUTH-EYE executes Android apps to trigger the OTP-relevant functionalities and then analyzes the OTP implementations including those proprietary ones. By only analyzing SMS responses, AUTH-EYE is able to assess the conformance of those implementations to our recommended rules and identify the potentially insecure apps. In our empirical study, AUTH-EYE analyzed 3,303 popular Android apps and found that 544 of them adopt SMS OTP authentication. The further analysis of AUTH-EYE demonstrated a far-from-optimistic status: the implementations of 536 (98.5%) out of the 544 apps violate at least one of our defined rules. The results indicate that Android app developers should seriously consider our discussed security rules and violations so as to implement SMS OTP properly.
Siqi Ma 0001, Runhan Feng, Juanru Li, Yang Liu 0118, Surya Nepal, Diethelm Ostry, Elisa Bertino, Robert H. Deng, Zhuo Ma 0001, Sanjay K. Jha
ACSAC4
2019 Privacy-Preserving Outsourced Speech Recognition for Smart IoT Devices
abstract
Most of the current intelligent Internet of Things (IoT) products take neural network-based speech recognition as the standard human–machine interaction interface. However, the traditional speech recognition frameworks for smart IoT devices always collect and transmit voice information in the form of plaintext, which may cause the disclosure of user privacy. Due to the wide utilization of speech features as biometric authentication, the privacy leakage can cause immeasurable losses to personal property and privacy. Therefore, in this paper, we propose an outsourced privacy-preserving speech recognition framework (OPSR) for smart IoT devices in the long short-term memory (LSTM) neural network and edge computing. In the framework, a series of additive secret sharing-based interactive protocols between two edge servers are designed to achieve lightweight outsourced computation. And based on the protocols, we implement the neural network training process of LSTM for intelligent IoT device voice control. Finally, combined with the universal composability theory and experiment results, we theoretically prove the correctness and security of our framework.
Zhuo Ma 0001, Yang Liu 0118, Ximeng Liu, Jianfeng Ma 0001, Feifei Li 0001
IEEE Internet Things J.2
2019 Lightweight Privacy-Preserving Ensemble Classification for Face Recognition
abstract
The development of machine learning technology and visual sensors is promoting the wider applications of face recognition into our daily life. However, if the face features in the servers are abused by the adversary, our privacy and wealth can be faced with great threat. Many security experts have pointed out that, by 3-D-printing technology, the adversary can utilize the leaked face feature data to masquerade others and break the E-bank accounts. Therefore, in this paper, we propose a lightweight privacy-preserving adaptive boosting (AdaBoost) classification framework for face recognition (POR) based on the additive secret sharing and edge computing. First, we improve the current additive secret sharing-based exponentiation and logarithm functions by expanding the effective input range. Then, by utilizing the protocols, two edge servers are deployed to cooperatively complete the ensemble classification of AdaBoost for face recognition. The application of edge computing ensures the efficiency and robustness of POR. Furthermore, we prove the correctness and security of our protocols by theoretic analysis. And experiment results show that, POR can reduce about 58% computation error compared with the existing differential privacy-based framework.
Zhuo Ma 0001, Yang Liu 0118, Ximeng Liu, Jianfeng Ma 0001, Kui Ren 0001
IEEE Internet Things J.2