Shuo Wang 0012

dblp:63/1591-12 · DBLP profile ↗
← Back
55ranked-venue papers
21as first author
43since 2021 · last 2026
0000-0001-8938-2364ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 26 · 6 first-author · 25 since 2021Artificial intelligence and machine learning · 14 · 7 first-author · 9 since 2021Databases, data management, data science and information retrieval · 7 · 4 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 4 first-author · 3 since 2021Software engineering, systems software and programming languages · 5 · 4 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Systems, architecture and hardware · 2 · 2 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 ExpShield: Safeguarding Web Text from Unauthorized Crawling and LLM Exploitation
Ruixuan Liu, Tianhao Wang 0001, Hongsheng Hu, Shuo Wang 0012, Li Xiong 0001
NDSS5
2026 SoK: Robustness in Large Language Models against Jailbreak Attacks
Feiyue Xu, Hongsheng Hu, Chaoxiang He, Sheng Hang, Hanqing Hu, Zhengyan Zhou, Bin B. Zhu, Shifeng Sun 0001, Dawu Gu, Shuo Wang 0012
SP12
2026 From Pixels to Trajectory: Universal Adversarial Example Detection via Temporal Imprints
abstract
We unveil discernible temporal (or historical) trajectory imprints resulting from adversarial example (AE) attacks. Standing in contrast to existing studies, which focus on spatial (or static) imprints within the targeted underlying victim models, we present a novel temporal paradigm for understanding these attacks. These imprints are encapsulated within a single loss metric, spanning universally across diverse tasks such as classification and regression, and modalities including image, text, and audio. Recognizing the distinct nature of loss between adversarial and clean examples, we exploit this temporal imprint for AE detection by proposing (Traceable Adversarial Temporal Imprints). TRAIT operates under minimal assumptions without prior knowledge of attacks, thereby framing the detection challenge as a one-class classification problem. However, detecting AEs is still challenged by significant overlaps between the constructed synthetic losses of adversarial and clean examples due to the absence of ground truth for incoming inputs. TRAIT addresses this challenge by converting the synthetic loss into a spectrum signature, using the technique of Fast Fourier Transform to highlight the discrepancies, drawing inspiration from the temporal nature of the imprints, analogous to time-series signals. Across 12 AE attacks including SMACK (USENIX Sec'2023), TRAIT demonstrates consistent outstanding performance across comprehensively evaluated modalities (image, text, audio), tasks (classification and regression), datasets (nine datasets), and model architectures (e.g., ResNeXt50, BERT, RoBERTa, AudioNet). In all scenarios, TRAIT achieves an AE detection accuracy exceeding 97%, often around 99%, while maintaining a false rejection rate of 1%. TRAIT remains effective under the formulated strong adaptive attacks.
Yansong Gao 0001, Huaibing Peng, Zhiyang Dai, Shuo Wang 0012, Hongsheng Hu, Anmin Fu, Minhui Xue 0001
IEEE Trans. Dependable Secur. Comput.5
2026 MaliVD: Source Code Vulnerability Localization via Attention-Based Multi-Modal Learning
abstract
Source code vulnerabilities represent a critical threat to software security, potentially leading to severe consequences such as data breaches and system failures. Traditional static analysis tools, while widely used, suffer from high false positive rates and struggle to adapt to the increasing complexity of modern software. Deep learning-based approaches hold promise for automated vulnerability detection, but they face challenges including limited dataset quality, inadequate feature extraction, and lack of precise vulnerability localization capabilities. To address these limitations, we propose MaliVD, a novel vulnerability detection method in source code that leverages a multi-modal attention mechanism. MaliVD not only identifies vulnerability types but also pinpoints the specific lines of code where vulnerabilities are triggered. The model extracts sequential, tree-based, and graph-based features from source code and employs specialized neural networks to learn these diverse representations. By strategically focusing on Points of Interest within the code, MaliVD effectively prioritizes potentially vulnerable code regions, enhancing both detection accuracy and localization precision. Experimental results show that when compared with eight advanced vulnerability detection models across three large datasets, MaliVD demonstrates superior vulnerability detection and localization capabilities, maintaining highF1scores and localization precision. Particularly on the ReliVul dataset, theF1score is improved by 21.82%, and the Top-5 localization accuracy is 18% higher than other methods, with lower false positives across six mainstream vulnerability types, validating MaliVD’s practical application value in real-world environments.
Enze Dai, Shuo Wang 0012, Xi Xiao 0001, Qing Li 0006, Sheng Wen, Tianqing Zhu
IEEE Trans. Inf. Forensics Secur.2
2026 DiffMI: Breaking Face Recognition Privacy via Diffusion-Driven Training-Free Model Inversion
abstract
Face recognition poses serious privacy risks due to its reliance on sensitive and immutable biometric data. While modern systems mitigate privacy risks by mapping facial images to embeddings (commonly regarded as privacy-preserving), model inversion attacks reveal that identity information can still be recovered, exposing critical vulnerabilities. However, existing attacks are often computationally expensive and lack generalization, especially those requiring target-specific training. Even training-free approaches suffer from limited identity controllability, hindering faithful reconstruction of nuanced or unseen identities. In this work, we propose DiffMI, the first diffusion-driven, training-free model inversion attack. DiffMI introduces a novel pipeline combining robust latent code initialization, a ranked adversarial refinement strategy, and a statistically grounded, confidence-aware optimization objective. DiffMI applies directly to unseen target identities and face recognition models, offering greater adaptability than training-dependent approaches while significantly reducing computational overhead. Our method achieves 84.42%–92.87% attack success rates against inversion-resilient systems and outperforms the best prior training-free GAN-based approach by 4.01%–9.82%. The implementation is available at https://github.com/azrealwang/DiffMI.
Hanrui Wang 0005, Shuo Wang 0012, Chun-Shien Lu, Isao Echizen
IEEE Trans. Inf. Forensics Secur.2
2025 LAMPS '25: ACM CCS Workshop on Large AI Systems and Models with Privacy and Security Analysis
abstract
With large AI systems and models (LAMs) playing an ever-growing role across diverse applications, their impact on the privacy and cybersecurity of critical infrastructure has become a pressing concern. The LAMPS workshop is dedicated to tackling these emerging challenges, promoting dialogue on cutting-edge developments and ethical issues in safeguarding LAMs within critical infrastructure contexts. Bringing together leading experts from around the world, this workshop will delve into the complex privacy and cybersecurity risks posed by LAMs in critical sectors. Attendees will explore innovative solutions, exchange best practices, and contribute to shaping the future research agenda, emphasizing the crucial balance between advancing AI technologies and securing critical digital and physical infrastructures.
Kwok-Yan Lam, Xiaoning Liu 0002, Derui Wang, Bo Li 0026, Wenyuan Xu 0001, Jieshan Chen, Minhui Xue 0001, Xingliang Yuan, Guangdong Bai, Shuo Wang 0012
CCS10
2025 Enhancing Adversarial Transferability with Checkpoints of a Single Model's Training
abstract
Adversarial attacks threaten the integrity of deep neural networks (DNNs), particularly in high-stakes applications. In this paper, we present a novel black-box adversarial attack that leverages the diverse checkpoints generated during a single model’s training trajectory. Unlike conventional ensemble attacks that require multiple surrogate models with diverse architectures, our approach exploits the intrinsic diversity captured over different training stages of a single surrogate model. By decomposing the learned representations into task-intrinsic and task-irrelevant components, we employ an accuracy gap-based selection strategy to identify checkpoints that predominantly capture transferable, task-intrinsic knowledge. Extensive experiments on ImageNet and CIFAR-10 demonstrate that our method consistently outperforms traditional ensemble attacks in terms of transferability, even under resource-constrained and practical settings. This work offers a resource-efficient solution for crafting highly transferable adversarial examples and provides new insights into the dynamics of adversarial vulnerability.
Shixin Li 0001, Chaoxiang He, Xiaojing Ma 0002, Bin B. Zhu, Shuo Wang 0012, Hongsheng Hu, Dongmei Zhang 0001, Linchen Yu
CVPR5
2025 Try to Poison My Deep Learning Data? Nowhere to Hide Your Trajectory Spectrum!
Yansong Gao 0001, Huaibing Peng, Zhi Zhang 0001, Shuo Wang 0012, Rayne Holland, Anmin Fu, Minhui Xue 0001, Derek Abbott
NDSS5
2025 BadFU: Backdoor Federated Learning through Adversarial Machine Unlearning
abstract
Federated learning (FL) has been widely adopted as a decentralized training paradigm that enables multiple clients to collaboratively learn a shared model without exposing their local data. As concerns over data privacy and regulatory compliance grow, machine unlearning, which aims to remove the influence of specific data from trained models, has become increasingly important in the federated setting to meet legal, ethical, or user-driven demands. However, integrating unlearning into FL introduces new challenges and raises largely unexplored security risks. In particular, adversaries may exploit the unlearning process to compromise the integrity of the global model. In this paper, we present the first backdoor attack in the context of federated unlearning, demonstrating that an adversary can inject backdoors into the global model through seemingly legitimate unlearning requests. Specifically, we propose BadFU, an attack strategy where a malicious client uses both backdoor and camouflage samples to train the global model normally during the federated training process. Once the client requests unlearning of the camouflage samples, the global model transitions into a backdoored state. Extensive experiments under various FL frameworks and unlearning strategies validate the effectiveness of BadFU, revealing a critical vulnerability in current federated unlearning practices and underscoring the urgent need for more secure and robust federated unlearning mechanisms.
Bingguang Lu, Hongsheng Hu, Yuantian Miao, Shaleeza Sohail, Chaoxiang He, Shuo Wang 0012, Xiao Chen 0002
RAID6
2025 Artificial intelligence security and privacy: a survey
abstract
Abstract Artificial intelligence (AI) is revolutionizing both industries and reshaping the global economy. However, the rapid advancement of AI technologies brings significant security and privacy challenges. Recent incidents highlight vulnerabilities in AI systems, such as data leakage and malicious code injection, leading to severe financial losses and privacy breaches. Although existing studies have discussed specific security threats, they often lack detailed granularity and cover a limited scope. In this survey, we fill this gap by systematically categorizing and analyzing the threats and countermeasures in AI systems, which span both the training and inference stages, encompass centralized and distributed settings, and address both conventional and foundation AI models. By reviewing existing literature, we aim to provide AI researchers and practitioners with a thorough understanding of system vulnerabilities and current countermeasures. We hope to inspire further research into robust solutions, ultimately contributing to the development of resilient AI technologies.
Xinlei He 0001, Guowen Xu, Xingshuo Han, Qian Wang 0002, Lingchen Zhao, Chao Shen 0001, Chenhao Lin, Zhengyu Zhao 0001, Qian Li 0024, Le Yang 0007, Shouling Ji, Shaofeng Li 0001, Haojin Zhu, Zhibo Wang 0001, Tianqing Zhu, Qi Li 0002, Chaoxiang He, Hongsheng Hu, Shuo Wang 0012, Shifeng Sun 0001, Hongwei Yao, Qinyu Zhang 0001, Kai Chen 0012, Yue Zhao 0027, Hongwei Li 0001, Xinyi Huang 0001, Dengguo Feng
Sci. China Inf. Sci.21
2025 Iterative Window Mean Filter: Thwarting Diffusion-Based Adversarial Purification
abstract
Face authentication systems have brought significant convenience and advanced developments, yet they have become unreliable due to their sensitivity to inconspicuous perturbations, such as adversarial attacks. Existing defenses often exhibit weaknesses when facing various attack algorithms and adaptive attacks or compromise accuracy for enhanced security. To address these challenges, we have developed a novel and highly efficient non-deep-learning-based image filter called the Iterative Window Mean Filter (IWMF) and proposed a new framework for adversarial purification, named IWMF-Diff, which integrates IWMF and denoising diffusion models. These methods can function as pre-processing modules to eliminate adversarial perturbations without necessitating further modifications or retraining of the target system. We demonstrate that our proposed methodologies fulfill four critical requirements: preserved accuracy, improved security, generalizability to various threats in different settings, and better resistance to adaptive attacks. This performance surpasses that of the state-of-the-art adversarial purification method, DiffPure. Our code is released athttps://github.com/azrealwang/iwmfdiff.
Hanrui Wang 0005, Ruoxi Sun 0001, Cunjian Chen, Minhui Xue 0001, Lay-Ki Soon, Shuo Wang 0012, Zhe Jin 0001
IEEE Trans. Dependable Secur. Comput.6
2025 RBLJAN: Robust Byte-Label Joint Attention Network for Network Traffic Classification
abstract
Network traffic classification plays a crucial role in network management and cyberspace security. As the Internet evolves with new applications and protocols, traditional machine learning-based methods relying on feature mining have become obsolete. Instead, deep learning-based methods are becoming more popular in the field of traffic classification due to their end-to-end processing approach. However, the vulnerability of neural networks to adversarial examples significantly compromises their performance. In this paper, we propose Robust Byte-Label Joint Attention Network (RBLJAN), an efficient and robust deep learning-based framework for encrypted network traffic classification at both the packet-level and the flow-level. RBLJAN comprises a classifier and an adversarial traffic generator. The classifier utilizes mechanisms such as header-payload parallel processing and byte-label joint attention learning to capture implicit correlations between bytes and labels, enabling the construction of powerful packet representations. The generator produces adversarial examples that are fed to the classifier to enhance its robustness. Experimental results demonstrate that RBLJAN achieves over 99% average F1-score on real-world legitimate traffic datasets and achieves 97.86% average F1-score on malware identification. Moreover, RBLJAN exhibits superior performance in terms of detection speed and robustness compared to state-of-the-art methods in real-world scenarios.
Xi Xiao 0001, Shuo Wang 0012, Guangwu Hu, Qing Li 0006, Kelong Mao, Xiapu Luo, Bin Zhang 0048, Shutao Xia
IEEE Trans. Dependable Secur. Comput.2
2024 LAMPS '24: ACM CCS Workshop on Large AI Systems and Models with Privacy and Safety Analysis
abstract
With large AI systems and models (LAMs) playing an ever-growing role across diverse applications, their impact on the privacy and cybersecurity of critical infrastructure has become a pressing concern. The LAMPS workshop is dedicated to tackling these emerging challenges, promoting dialogue on cutting-edge developments and ethical issues in safeguarding LAMs within critical infrastructure contexts. Bringing together leading experts from around the world, this workshop will delve into the complex privacy and cybersecurity risks posed by LAMs in critical sectors. Attendees will explore innovative solutions, exchange best practices, and contribute to shaping the future research agenda, emphasizing the crucial balance between advancing AI technologies and securing critical digital and physical infrastructures.
Bo Li 0026, Wenyuan Xu 0001, Jieshan Chen, Yang Zhang 0016, Minhui Xue 0001, Shuo Wang 0012, Guangdong Bai, Xingliang Yuan
CCS6
2024 Learning with Mixture of Prototypes for Out-of-Distribution Detection
abstract
Out-of-distribution (OOD) detection aims to detect testing samples far away from the in-distribution (ID) training data, which is crucial for the safe deployment of machine learning models in the real world. Distance-based OOD detection methods have emerged with enhanced deep representation learning. They identify unseen OOD samples by measuring their distances from ID class centroids or prototypes. However, existing approaches learn the representation relying on oversimplified data assumptions, e.g. modeling ID data of each class with one centroid class prototype or using loss functions not designed for OOD detection, which overlook the natural diversities within the data. Naively enforcing data samples of each class to be compact around only one prototype leads to inadequate modeling of realistic data and limited performance. To tackle these issues, we propose PrototypicAl Learning with a Mixture of prototypes (PALM) that models each class with multiple prototypes to capture the sample diversities, which learns more faithful and compact samples embeddings for enhanching OOD detection. Our method automatically identifies and dynamically updates prototypes, assigning each sample to a subset of prototypes via reciprocal neighbor soft assignment weights. To learn embeddings with multiple prototypes, PALM optimizes a maximum likelihood estimation (MLE) loss to encourage the sample embeddings to compact around the associated prototypes, as well as a contrastive loss on all prototypes to enhance intra-class compactness and inter-class discrimination at the prototype level. Compared to previous methods with prototypes, the proposed mixture prototype modeling of PALM promotes the representations of each ID class to be more compact and separable from others and the unseen OOD samples, resulting in more reliable OOD detection. Moreover, the automatic estimation of prototypes enables our approach to be extended to the challenging OOD detection task with unlabelled ID data. Extensive experiments demonstrate the superiority of PALM over previous methods, achieving state-of-the-art average AUROC performance of 93.82 on the challenging CIFAR-100 benchmark.
Haodong Lu 0002, Dong Gong, Shuo Wang 0012, Minhui Xue 0001, Lina Yao 0001, Kristen Moore
ICLR3
2024 A Duty to Forget, a Right to be Assured? Exposing Vulnerabilities in Machine Unlearning Services
Hongsheng Hu, Shuo Wang 0012, Jiamin Chang, Haonan Zhong, Ruoxi Sun 0001, Shuang Hao 0001, Haojin Zhu, Minhui Xue 0001
NDSS2
2024 GraphGuard: Detecting and Counteracting Training Data Misuse in Graph Neural Networks
Bang Wu 0004, He Zhang 0012, Xiangwen Yang, Shuo Wang 0012, Minhui Xue 0001, Shirui Pan, Xingliang Yuan
NDSS4
2024 Learn What You Want to Unlearn: Unlearning Inversion Attacks against Machine Unlearning
abstract
Machine unlearning has become a promising solution for fulfilling the "right to be forgotten", under which individuals can request the deletion of their data from machine learning models. However, existing studies of machine unlearning mainly focus on the efficacy and efficiency of unlearning methods, while neglecting the investigation of the privacy vulnerability during the unlearning process. With two versions of a model available to an adversary, that is, the original model and the unlearned model, machine unlearning opens up a new attack surface. In this paper, we conduct the first investigation to understand the extent to which machine unlearning can leak the confidential content of the unlearned data. Specifically, under the Machine Learning as a Service setting, we propose unlearning inversion attacks that can reveal the feature and label information of an unlearned sample by only accessing the original and unlearned model. The effectiveness of the proposed unlearning inversion attacks is evaluated through extensive experiments on benchmark datasets across various model architectures and on both exact and approximate representative unlearning approaches. The experimental results indicate that the proposed attack can reveal the sensitive information of the unlearned data. As such, we identify three possible defenses that help to mitigate the proposed attacks, while at the cost of reducing the utility of the unlearned model. The study in this paper uncovers an underexplored gap between machine unlearning and the privacy of unlearned data, highlighting the need for the careful design of mechanisms for implementing unlearning without leaking the information of the unlearned data.
Hongsheng Hu, Shuo Wang 0012, Tian Dong 0003, Minhui Xue 0001
SP2
2024 LACMUS: Latent Concept Masking for General Robustness Enhancement of DNNs
abstract
The susceptibility of Deep Neural Networks (DNNs) to adversarial attacks and their limited robustness to real-world variations pose substantial challenges to their widespread adoption. Adversarial training has shown promise in fortifying models against such perturbations, however current methods are often specific to a single type of attack and can significantly diminish the model’s overall performance. In response, we present LAtent Concept Masking for robUStness (LACMUS), a novel perceptually-driven methodology that enhances DNN robustness without requiring prior knowledge about the adversarial contexts. We argue that DNNs’ sensitivity to adversarial perturbations and distribution drifts stems from overfitting to non-common concepts within the dataset, leading to an over-reliance on specific learned instances and increased vulnerability. LACMUS addresses this by mapping high-dimensional data into a latent conceptual space to identify and navigate patterns of "non-common concepts" within the latent concept space. It then applies a concept masking strategy to selectively obscure data features, prompting the model to base its decisions on a wider array of information and thus enhancing its decision-making robustness. LACMUS distinguishes itself as a versatile, attack-agnostic framework that employs concept-wise augmentation to enhance robustness against a spectrum of adversarial, semantic, and distributional challenges. Our contributions include the development of a tool for robustness enhancement, a mechanism for mapping data to latent concept space, a strategy for identifying patterns of concept-wise misclassification, and a novel data augmentation module that leverages latent concepts. LACMUS is proven to enhance model resilience and generalization, even when training data is scarce, with experiments on MNIST, CIFAR-10, ImageNet, and CelebA supporting its effectiveness. We also provide augmented datasets to the research community, bolstering the robustness of models trained on them.
Shuo Wang 0012, Hongsheng Hu, Jiamin Chang, Benjamin Zi Hao Zhao, Minhui Xue 0001
SP1
2024 Securing Graph Neural Networks in MLaaS: A Comprehensive Realization of Query-based Integrity Verification
abstract
The deployment of Graph Neural Networks (GNNs) within Machine Learning as a Service (MLaaS) has opened up new attack surfaces and an escalation in security concerns regarding model-centric attacks. These attacks can directly manipulate the GNN model parameters during serving, causing incorrect predictions and posing substantial threats to essential GNN applications. Traditional integrity verification methods falter in this context due to the limitations imposed by MLaaS and the distinct characteristics of GNN models.In this research, we introduce a groundbreaking approach to protect GNN models in MLaaS from model-centric attacks. Our approach includes a comprehensive verification schema for GNN’s integrity, taking into account both transductive and inductive GNNs, and accommodating varying pre-deployment knowledge of the models. We propose a query-based verification technique, fortified with innovative node fingerprint generation algorithms. To deal with advanced attackers who know our mechanisms in advance, we introduce randomized fingerprint nodes within our design. The experimental evaluation demonstrates that our method can detect five representative adversarial model-centric attacks, displaying 2 to 4 times greater efficiency compared to baselines.
Bang Wu 0004, Xingliang Yuan, Shuo Wang 0012, Qi Li 0002, Minhui Xue 0001, Shirui Pan
SP3
2024 DNN-GP: Diagnosing and Mitigating Model's Faults Using Latent Concepts
Shuo Wang 0012, Hongsheng Hu, Jiamin Chang, Benjamin Zi Hao Zhao, Qi Alfred Chen, Minhui Xue 0001
USENIX Security Symposium1
2024 Cardinality Counting in "Alcatraz": A Privacy-aware Federated Learning Approach
abstract
The task of cardinality counting, pivotal for data analysis, endeavors to quantify unique elements within datasets and has significant applications across various sectors like healthcare, marketing, cybersecurity, and web analytics. Current methods, categorized into deterministic and probabilistic, often fail to prioritize data privacy. Given the fragmentation of datasets across various organizations, there is an elevated risk of inadvertently disclosing sensitive information during collaborative data studies using state-of-the-art cardinality counting techniques. This study introduces an innovative privacy-centric solution for the cardinality counting dilemma, leveraging a federated learning framework. Our approach involves employing a locally differentially private data encoding for initial processing, followed by a privacy-aware federated K-means clustering strategy, ensuring that cardinality counting occurs across distinct datasets without necessitating data amalgamation. The efficacy of our methodology is underscored by promising results from tests on both real-world and simulated datasets, pointing towards a transformative approach to privacy-sensitive cardinality counting in contemporary data science.
Nan Wu 0013, Xin Yuan 0004, Shuo Wang 0012, Hongsheng Hu, Minhui Xue 0001
WWW3
2024 SoK: Can Trajectory Generation Combine Privacy and Utility?
abstract
While location trajectories represent a valuable data source for analyses and location-based services, they can reveal sensitive information, such as political and religious preferences. Differentially private publication mechanisms have been proposed to allow for analyses under rigorous privacy guarantees. However, the traditional protection schemes suffer from a limiting privacy-utility trade-off and are vulnerable to correlation and reconstruction attacks. Synthetic trajectory data generation and release represent a promising alternative to protection algorithms. While initial proposals achieve remarkable utility, they fail to provide rigorous privacy guarantees. This paper proposes a framework for designing a privacy-preserving trajectory publication approach by defining five design goals, particularly stressing the importance of choosing an appropriate Unit of Privacy. Based on this framework, we briefly discuss the existing trajectory protection approaches, emphasising their shortcomings. This work focuses on the systematisation of the state-of-the-art generative models for trajectories in the context of the proposed framework. We find that no existing solution satisfies all requirements. Thus, we perform an experimental study evaluating the applicability of six sequential generative models to the trajectory domain. Finally, we conclude that a generative trajectory model providing semantic guarantees remains an open research question and propose concrete next steps for future research.
Erik Buchholz, Alsharif Abuadbba, Shuo Wang 0012, Surya Nepal, Salil S. Kanhere
Proc. Priv. Enhancing Technol.3
2024 On Model Outsourcing Adaptive Attacks to Deep Learning Backdoor Defenses
abstract
Deep learning models with backdoors act maliciously when triggered but seem normal otherwise. This risk, often increased by model outsourcing, challenges their secure use. Although countermeasures exist, their defense against adaptive attacks is under-examined, possibly leading to security misjudgments. This study is the first intricate examination illustrating the difficulty of detecting backdoors in outsourced models, especially when attackers adjust their strategies, even if their capabilities are significantly limited. It is relatively straightforward for attackers to circumvent detection by trivially violating its threat model (e.g., using advanced backdoor types or trigger designs not covered by the detection). However, this research highlights that various leading detection defenses can simultaneously be evaded using simple adaptive strategies, even under their defined threat models and with limited adversary capabilities (e.g., using easily detectable triggers while maintaining a high attack success rate). To be more specific, this study introduces a novel methodology that employs trigger specificity enhancement and training regulation in a symbiotic manner. This approach allows us to evade multiple backdoor detection defenses simultaneously, including Neural Cleanse (Oakland 19’), ABS (CCS 19’), and MNTD (Oakland 21’). These were the detection tools selected for the Evasive Trojans Track of the 2022 NeurIPS Trojan Detection Challenge. Even when applied in conjunction with these defenses under stringent conditions, such as a high attack success rate (> 97%) and the restricted use of the simplest trigger (small white square), our straightforward method garnered the second prize in NeurIPS Trojan Detection Challenge. Notably, for the first time, our adaptive attack successfully evaded other recent state-of-the-art defenses, including FeatureRE (NeurIPS 22’) and Beatrix (NDSS 23’). This study suggests that existing model outsourcing backdoor defenses remain vulnerable to adaptive attacks, and thus, the use of third-party models should be avoided whenever possible.
Huaibing Peng, Huming Qiu, Shuo Wang 0012, Anmin Fu, Said F. Al-Sarawi, Derek Abbott, Yansong Gao 0001
IEEE Trans. Inf. Forensics Secur.4
2024 Generating Semantic Adversarial Examples via Feature Manipulation in Latent Space
abstract
The susceptibility of deep neural networks (DNNs) to adversarial intrusions, exemplified by adversarial examples, is well-documented. Conventional attacks implement unstructured, pixel-wise perturbations to mislead classifiers, which often results in a noticeable departure from natural samples and lacks human-perceptible interpretability. In this work, we present an adversarial attack strategy that implements fine-granularity, semantic-meaning-oriented structural perturbations. Our proposed methodology manipulates the semantic attributes of images through the use of disentangled latent codes. We engineer adversarial perturbations by manipulating either a single latent code or a combination thereof. To this end, we propose two unsupervised semantic manipulation strategies: one based on vector-disentangled representation and the other on feature map-disentangled representation, taking into consideration the complexity of the latent codes and the smoothness of the reconstructed images. Our empirical evaluations, conducted extensively on real-world image data, showcase the potency of our attacks, particularly against black-box classifiers. Furthermore, we establish the existence of a universal semantic adversarial example that is agnostic to specific images.
Shuo Wang 0012, Shangyu Chen, Surya Nepal, Carsten Rudolph, Marthie Grobler
IEEE Trans. Neural Networks Learn. Syst.1
2024 A Multi-Task Adversarial Attack against Face Authentication
abstract
Deep learning-based identity management systems, such as face authentication systems, are vulnerable to adversarial attacks. However, existing attacks are typically designed for single-task purposes, which means they are tailored to exploit vulnerabilities unique to the individual target rather than being adaptable for multiple users or systems. This limitation makes them unsuitable for certain attack scenarios, such as morphing, universal, transferable, and counterattacks. In this article, we propose a multi-task adversarial attack algorithm called MTADV that are adaptable for multiple users or systems. By interpreting these scenarios as multi-task attacks, MTADV is applicable to both single- and multi-task attacks, and feasible in the white- and gray-box settings. Furthermore, MTADV is effective against various face datasets, including LFW, CelebA, and CelebA-HQ, and can work with different deep learning models, such as FaceNet, InsightFace, and CurricularFace. Importantly, MTADV retains its feasibility as a single-task attack targeting a single user/system. To the best of our knowledge, MTADV is the first adversarial attack method that can target all of the aforementioned scenarios in one algorithm.
Hanrui Wang 0005, Shuo Wang 0012, Cunjian Chen, Massimo Tistarelli, Zhe Jin 0001
ACM Trans. Multim. Comput. Commun. Appl.2
2023 DeepTaster: Adversarial Perturbation-Based Fingerprinting to Identify Proprietary Dataset Use in Deep Neural Networks
abstract
Training deep neural networks (DNNs) requires large datasets and powerful computing resources, which has led some owners to restrict redistribution without permission. Watermarking techniques that embed confidential data into DNNs have been used to protect ownership, but these can degrade model performance and are vulnerable to watermark removal attacks. Recently, DeepJudge was introduced as an alternative approach to measuring the similarity between a suspect and a victim model. While DeepJudge shows promise in addressing the shortcomings of watermarking, it primarily addresses situations where the suspect model copies the victim’s architecture. In this study, we introduce DeepTaster, a novel DNN fingerprinting technique, to address scenarios where a victim’s data is unlawfully used to build a suspect model. DeepTaster can effectively identify such DNN model theft attacks, even when the suspect model’s architecture deviates from the victim’s. To accomplish this, DeepTaster generates adversarial images with perturbations, transforms them into the Fourier frequency domain, and uses these transformed images to identify the dataset used in a suspect model. The underlying premise is that adversarial images can capture the unique characteristics of DNNs built with a specific dataset. To demonstrate the effectiveness of DeepTaster, we evaluated the effectiveness of DeepTaster by assessing its detection accuracy on three datasets (CIFAR10, MNIST, and Tiny-ImageNet) across three model architectures (ResNet18, VGG16, and DenseNet161). We conducted experiments under various attack scenarios, including transfer learning, pruning, fine-tuning, and data augmentation. Specifically, in the Multi-Architecture Attack scenario, DeepTaster was able to identify all the stolen cases across all datasets, while DeepJudge failed to detect any of the cases.
Seonhye Park, Alsharif Abuadbba, Shuo Wang 0012, Kristen Moore, Yansong Gao 0001, Hyoungshick Kim, Surya Nepal
ACSAC3
2023 Demystifying Uneven Vulnerability of Link Stealing Attacks against Graph Neural Networks
abstract
While graph neural networks (GNNs) dominate the state-of-the-art for exploring graphs in real-world applications, they have been shown to be vulnerable to a growing number of privacy attacks. For instance, link stealing is a well-known membership inference attack (MIA) on edges that infers the presence of an edge in a GNN’s training graph. Recent studies on independent and identically distributed data (e.g., images) have empirically demonstrated that individuals from different groups suffer from different levels of privacy risks to MIAs, i.e., uneven vulnerability. However, theoretical evidence of such uneven vulnerability is missing. In this paper, we first present theoretical evidence of the uneven vulnerability of GNNs to link stealing attacks, which lays the foundation for demystifying such uneven risks among different groups of edges. We further demonstrate a group-based attack paradigm to expose the practical privacy harm to GNN users derived from the uneven vulnerability of edges. Finally, we empirically validate the existence of obvious uneven vulnerability on nine real-world datasets (e.g., about 25% AUC difference between different groups in the Credit graph). Compared with existing methods, the outperformance of our group-based attack paradigm confirms that customising different strategies for different groups results in more effective privacy attacks.
He Zhang 0012, Bang Wu 0004, Shuo Wang 0012, Xiangwen Yang, Minhui Xue 0001, Shirui Pan, Xingliang Yuan
ICML3
2023 DOITRUST: Dissecting On-chain Compromised Internet Domains via Graph Learning
Shuo Wang 0012, Mahathir Almashor, Alsharif Abuadbba, Ruoxi Sun 0001, Minhui Xue 0001, Calvin Wang, Raj Gaire 0001, Surya Nepal, Seyit Ahmet Çamtepe
NDSS1
2023 Mate! Are You Really Aware? An Explainability-Guided Testing Framework for Robustness of Malware Detectors
abstract
Numerous open-source and commercial malware detectors are available. However, their efficacy is threatened by new adversarial attacks, whereby malware attempts to evade detection, e.g., by performing feature-space manipulation. In this work, we propose an explainability-guided and model-agnostic testing framework for robustness of malware detectors when confronted with adversarial attacks. The framework introduces the concept of Accrued Malicious Magnitude (AMM) to identify which malware features could be manipulated to maximize the likelihood of evading detection. We then use this framework to test several state-of-the-art malware detectors' ability to detect manipulated malware. We find that (i) commercial antivirus engines are vulnerable to AMM-guided test cases; (ii) the ability of a manipulated malware generated using one detector to evade detection by another detector (i.e., transferability) depends on the overlap of features with large AMM values between the different detectors; and (iii) AMM values effectively measure the fragility of features (i.e., capability of feature-space manipulation to flip the prediction results) and explain the robustness of malware detectors facing evasion attacks. Our findings shed light on the limitations of current malware detectors, as well as how they can be improved.
Ruoxi Sun 0001, Minhui Xue 0001, Gareth Tyson, Tian Dong 0003, Shaofeng Li 0001, Shuo Wang 0012, Haojin Zhu, Seyit Ahmet Çamtepe, Surya Nepal
ESEC/SIGSOFT FSE6
2023 PublicCheck: Public Integrity Verification for Services of Run-time Deep Models
abstract
Existing integrity verification approaches for deep models are designed for private verification (i.e., assuming the service provider is honest, with white-box access to model parameters). However, private verification approaches do not allow model users to verify the model at run-time. Instead, they must trust the service provider, who may tamper with the verification results. In contrast, a public verification approach that considers the possibility of dishonest service providers can benefit a wider range of users. In this paper, we propose PublicCheck, a practical public integrity verification solution for services of run-time deep models. PublicCheck considers dishonest service providers, and overcomes public verification challenges of being lightweight, providing anti-counterfeiting protection, and having fingerprinting samples that appear smooth. To capture and fingerprint the inherent prediction behaviors of a run-time model, PublicCheck generates smoothly transformed and augmented encysted samples that are enclosed around the model's decision boundary while ensuring that the verification queries are indistinguishable from normal queries. PublicCheck is also applicable when knowledge of the target model is limited (e.g., with no knowledge of gradients or model parameters). A thorough evaluation of PublicCheck demonstrates the strong capability for model integrity breach detection (100% detection accuracy with less than 10 black-box API queries) against various model integrity attacks and model compression attacks. PublicCheck also demonstrates the smooth appearance, feasibility, and efficiency of generating a plethora of encysted samples for fingerprinting.
Shuo Wang 0012, Alsharif Abuadbba, Sidharth Agarwal, Kristen Moore, Ruoxi Sun 0001, Minhui Xue 0001, Surya Nepal, Seyit Ahmet Çamtepe, Salil S. Kanhere
SP1
2023 Not Seen, Not Heard in the Digital World! Measuring Privacy Practices in Children's Apps
abstract
The digital age has brought a world of opportunity to children. Connectivity can be a game-changer for some of the world’s most marginalized children. However, while legislatures around the world have enacted regulations to protect children’s online privacy, and app stores have instituted various protections, privacy in mobile apps remains a growing concern for parents and wider society. In this paper, we explore the potential privacy issues and threats that exist in these apps. We investigate 20195 mobile apps from the Google Play store that are designed particularly for children (Family apps) or include children in their target user groups (Normal apps). Using both static and dynamic analysis, we find that 4.47% of Family apps request location permissions, even though collecting location information from children is forbidden by the Play store, and 81.25% of Family apps use trackers (which are not allowed in children’s apps). Even major developers with 40+ kids apps on the Play store use ad trackers. Furthermore, we find that most permission request notifications are not well designed for children, and 19.25% apps have inconsistent content age ratings across the different protection authorities. Our findings suggest that, despite significant attention to children’s privacy, a large gap between regulatory provisions, app store policies, and actual development practices exist. Our research sheds light for government policymakers, app stores, and developers.
Ruoxi Sun 0001, Minhui Xue 0001, Gareth Tyson, Shuo Wang 0012, Seyit Ahmet Çamtepe, Surya Nepal
WWW4
2023 Text classification on heterogeneous information network via enhanced GCN and knowledge
Shuo Wang 0012, Yunpeng Cui
Neural Comput. Appl.3
2023 Defeating Misclassification Attacks Against Transfer Learning
abstract
Transfer learning is prevalent as a technique to efficiently generate new models (Student models) based on the knowledge transferred from a pre-trained model (Teacher model). However, Teacher models are often publicly available for sharing and reuse, which inevitably introduces vulnerability to trigger severe attacks against transfer learning systems. In this article, we take a first step towards mitigating one of the most advanced misclassification attacks in transfer learning. We design a distilleddifferentiatorvia activation-based network pruning to enervate the attack transferability while retaining accuracy. We adopt an ensemble structure from variant differentiators to improve the defence robustness. To avoid the bloated ensemble size during inference, we propose a two-phase defence, in which inference from the Student model is first performed to narrow down the candidate differentiators to be assembled, and later only a small, fixed number of them can be chosen to validate clean or reject adversarial inputs effectively. Our comprehensive evaluations on both large and small image recognition tasks confirm that the Student models with our defence of only 5 differentiators are immune to over 90% of the adversarial inputs with an accuracy loss of less than 10%. Our comparison also demonstrates that our design outperforms prior problematic defences.
Bang Wu 0004, Shuo Wang 0012, Xingliang Yuan, Cong Wang 0001, Carsten Rudolph, Xiangwen Yang
IEEE Trans. Dependable Secur. Comput.2
2022 Reconstruction Attack on Differential Private Trajectory Protection Mechanisms
abstract
Location trajectories collected by smartphones and other devices represent a valuable data source for applications such as location-based services. Likewise, trajectories have the potential to reveal sensitive information about individuals, e.g., religious beliefs or sexual orientations. Accordingly, trajectory datasets require appropriate sanitization. Due to their strong theoretical privacy guarantees, differential private publication mechanisms receive much attention. However, the large amount of noise required to achieve differential privacy yields structural differences, e.g., ship trajectories passing over land. We propose a deep learning-based Reconstruction Attack on Protected Trajectories (RAoPT), that leverages the mentioned differences to partly reconstruct the original trajectory from a differential private release. The evaluation shows that our RAoPT model can reduce the Euclidean and Hausdorff distances between the released and original trajectories by over 68 % on two real-world datasets under protection with ε ≤ 1. In this setting, the attack increases the average Jaccard index of the trajectories’ convex hulls, representing a user’s activity space, by over 180 %. Trained on the GeoLife dataset, the model still reduces the Euclidean and Hausdorff distances by over 60 % for T-Drive trajectories protected with a state-of-the-art mechanism (ε = 0.1). This work highlights shortcomings of current trajectory publication mechanisms, and thus motivates further research on privacy-preserving publication schemes.
Erik Buchholz, Alsharif Abuadbba, Shuo Wang 0012, Surya Nepal, Salil S. Kanhere
ACSAC3
2022 Latent Space-Based Backdoor Attacks Against Deep Neural Networks
abstract
The outstanding performance of modern deep learning systems resulted in their widespread adoption in various application domains, which include security-critical applications. However, recent works have shown that these systems are vulnerable to backdoor attacks. This paper proposed a novel approach to perform latent backdoor attacks. Instead of designing the exogenetic trigger backdoor on the pixel space, which has been done by existing works, this paper explored the connection between latent space manipulation and endogenic backdoor trigger generation by utilising deep generative models to generate the backdoor trigger in the latent space. The effectiveness of the proposed attack is demonstrated on several neural network architectures trained on three well-known datasets, which are MNIST, CIFAR-10 and GTSRB. This study is undertaken to provide a new viewpoint for better understanding the endogenic vulnerability of the deep neural networks due to the lack of training data and test data, instead of creating new exogenetic misclassification behaviours for existing backdoor attacks.
Adrian Kristanto, Shuo Wang 0012, Carsten Rudolph
IJCNN2
2022 R-Net: Robustness Enhanced Financial Time-Series Prediction with Differential Privacy
abstract
Artificial intelligence has been investigated to conduct automatic predictions on financial time series such as stock. However, they faced two challenges. Firstly, the stock movement is affected by both technical fundamentals and external textual information. Secondly, resource data are often highly-noisy and heterogeneous, and prediction based on noisy data usually leads to significant errors. We propose a robust and accurate model (R-Net) that incorporates both technical information and qualitative sentiment derived from news reports for daily stock movement prediction in response to these challenges. A variety of enhancement strategies are adopted to improve the prediction model's robustness and accuracy. Specifically, a multimodal CNN and LSTM neural networks are applied to extract semantics from text and model complex temporal characteristics for stock market prediction. Further, based on the connection between the robustness of deep neural networks and differential privacy, we utilize provable noise injection and heterogeneous Gaussian mechanisms to enhance model robustness and accuracy. Experimental results on S&P 500 stocks demonstrate that our proposed R-Net, which integrates four enhancements, achieves 12.7% and 0.67% improvement in prediction accuracy for trends and price value prediction, respectively.
Shuo Wang 0012, Jinyuan Qin, Carsten Rudolph, Surya Nepal, Marthie Grobler
IJCNN1
2022 Adversarial Detection by Latent Style Transformations
abstract
Detection-based defense approaches are effective against adversarial attacks without compromising the structure of the protected model. However, they could be bypassed by stronger adversarial attacks and are limited in their ability to handle high-fidelity images. In this paper, we explore an effective detection-based defense against adversarial attacks on images (including high-resolution images) by extending the investigation beyond a single-instance perspective to incorporate its transformations as well. Our intuition is that the essential characteristics of a valid image are generally not affected by non-essential style transformations, for example, a slight variation in the facial expression of a portrait would not alter its identification. In contrast, adversarial examples are designed to affect only a single instance at a time, with unpredictable effects on a set of transformations of the instance. Consequently, we leverage a controllable generative mechanism to conduct the non-essential style transformations for a given image via modification along the style axis in the latent space. Next, the consistency of prediction between the given input and its style transformations is used to distinguish adversarial instances. Based on experiments on three image datasets, including high-resolution images, we demonstrated that our defense could detect 90–100 percent of adversarial examples produced by various state-of-the-art adversarial attacks, with a low false-positive rate.
Shuo Wang 0012, Surya Nepal, Alsharif Abuadbba, Carsten Rudolph, Marthie Grobler
IEEE Trans. Inf. Forensics Secur.1
2022 OCTOPUS: Overcoming Performance and Privatization Bottlenecks in Distributed Learning
abstract
The diversity and quantity of data warehouses, gathering data from distributed devices such as mobile devices, can enhance the success and robustness of machine learning algorithms. Federated learning enables distributed participants to collaboratively learn a commonly shared model while holding data locally. However, it is also faced with expensive communication and limitations due to the heterogeneity of distributed data sources and lack of access to global data. In this paper, we investigate a practical distributed learning scenario where multiple downstream tasks (e.g., classifiers) could be efficiently learned from dynamically updated and non-iid distributed data sources while providing local data privatization. We introduce a new distributed/collaborative learning scheme to address communication overhead via latent compression, leveraging global data while providing privatization of local data without additional cost due to encryption or perturbation. This scheme divides learning into (1) informative feature encoding, and transmitting the latent representation of local data to address communication overhead; (2) downstream tasks centralized at the server using the encoded codes gathered from each node to address computing overhead. Besides, a disentanglement strategy is applied to address the privatization of sensitive components of local data. Extensive experiments are conducted on image and speech datasets. The results demonstrate that downstream tasks with the compact latent representations with the privatization of local data can achieve comparable accuracy to centralized learning.
Shuo Wang 0012, Surya Nepal, Kristen Moore, Marthie Grobler, Carsten Rudolph, Alsharif Abuadbba
IEEE Trans. Parallel Distributed Syst.1
2022 Backdoor Attacks Against Transfer Learning With Pre-Trained Deep Learning Models
abstract
Transfer learning provides an effective solution for feasibly and fast customize accurateStudentmodels, by transferring the learned knowledge of pre-trainedTeachermodels over large datasets via fine-tuning. Many pre-trained Teacher models used in transfer learning are publicly available and maintained by public platforms, increasing their vulnerability to backdoor attacks. In this article, we demonstrate a backdoor threat to transfer learning tasks on both image and time-series data leveraging the knowledge of publicly accessible Teacher models, aimed at defeating three commonly adopted defenses:pruning-based,retraining-basedandinput pre-processing-based defenses. Specifically, ($\mathcal {A}$A) ranking-based selection mechanism to speed up the backdoor trigger generation and perturbation process while defeatingpruning-basedand/orretraining-based defenses. ($\mathcal {B}$B) autoencoder-powered trigger generation is proposed to produce a robust trigger that can defeat theinput pre-processing-based defense, while guaranteeing that selected neuron(s) can be significantly activated. ($\mathcal {C}$C) defense-aware retraining to generate the manipulated model using reverse-engineered model inputs. We launch effective misclassification attacks on Student models over real-world images, brain Magnetic Resonance Imaging (MRI) data and Electrocardiography (ECG) learning systems. The experiments reveal that our enhanced attack can maintain the 98.4 and 97.2 percent classification accuracy as the genuine model on clean image and time series inputs while improving$27.9\%-100\%$27.9%-100%and$27.1\%-56.1\%$27.1%-56.1%attack success rate on trojaned image and time series inputs respectively in the presence of pruning-based and/or retraining-based defenses.
Shuo Wang 0012, Surya Nepal, Carsten Rudolph, Marthie Grobler, Shangyu Chen
IEEE Trans. Serv. Comput.1
2022 Defending Adversarial Attacks via Semantic Feature Manipulation
abstract
Machine learning models have demonstrated vulnerability to adversarial attacks, more specifically misclassification of adversarial examples. In this article, we propose a one-off and attack-agnostic Feature Manipulation (FM)-Defense to detect and purify adversarial examples in an interpretable and efficient manner. The intuition is that the classification result of a normal image is generally resistant to non-significant intrinsic feature changes, e.g., varying the thickness of handwritten digits. In contrast, adversarial examples are sensitive to such changes since the perturbation lacks transferability. To enable manipulation of features, a Combo-variational autoencoder is applied to learn disentangled latent codes that reveal semantic features. The resistance to classification change over the morphs, derived by varying and reconstructing latent codes, is used to detect suspicious inputs. Furthermore, Combo-VAE is enhanced to purify the adversarial examples with good quality by considering class-shared and class-unique features. We empirically demonstrate the effectiveness of detection and quality of purified instances. Our experiments on three datasets show that FM-Defense can detect nearly 100 percent of adversarial examples produced by different state-of-the-art adversarial attacks. It achieves more than 99 percent overall purification accuracy on the suspicious instances that close the manifold of clean examples.
Shuo Wang 0012, Surya Nepal, Carsten Rudolph, Marthie Grobler, Shangyu Chen, Zike An
IEEE Trans. Serv. Comput.1
2021 Projective Ranking: A Transferable Evasion Attack Method on Graph Neural Networks
abstract
Graph Neural Networks (GNNs) have emerged as a series of effective learning methods for graph-related tasks. However, GNNs are shown vulnerable to adversarial attacks, where attackers can fool GNNs into making wrong predictions on adversarial samples with well-designed perturbations. Specifically, we observe that the current evasion attacks suffer from two limitations: (1) the attack strategy based on the reinforcement learning method might not be transferable when the attack budget changes; (2) the greedy mechanism in the vanilla gradient-based method ignores the long-term benefits of each perturbation operation. In this paper, we propose a new attack method named projective ranking to overcome the above limitations. Our idea is to learn a powerful attack strategy considering the long-term benefits of perturbations, then adjust it as little as possible to generate adversarial samples under different budgets. We further employ mutual information to measure the long-term benefits of each perturbation and rank them accordingly, so the learned attack strategy has better attack performance. Our method dramatically reduces the adaptation cost of learning a new attack strategy by projecting the attack strategy when the attack budget changes. Our preliminary evaluation results in synthesized and real-world datasets demonstrate that our method owns powerful attack performance and effective transferability.
He Zhang 0012, Bang Wu 0004, Xiangwen Yang, Chuan Zhou 0001, Shuo Wang 0012, Xingliang Yuan, Shirui Pan
CIKM5
2021 Similarity-based Gray-box Adversarial Attack Against Deep Face Recognition
abstract
The majority of adversarial attack techniques perform well against deep face recognition when the full knowledge of the system is revealed (white-box). However, such techniques act unsuccessfully in the gray-box setting where the face templates are unknown to the attackers. In this work, we propose a similarity-based gray-box adversarial attack (SGADV) technique with a newly developed objective function. SGADV utilizes the dissimilarity score to produce the optimized adversarial example, i.e., similarity-based adversarial attack. This technique applies to both white-box and gray-box attacks against authentication systems that determine genuine or imposter users using the dissimilarity score. To validate the effectiveness of SGADV, we conduct extensive experiments on face datasets of LFW, CelebA, and CelebA-HQ against deep face recognition models of FaceNet and InsightFace in both white-box and gray-box settings. The results suggest that the proposed method significantly outperforms the existing adversarial attack techniques in the gray-box setting. We hence summarize that the similarity-base approaches to develop the adversarial example could satisfactorily cater to the gray-box attack scenarios for de-authentication.
Hanrui Wang 0003, Shuo Wang 0012, Zhe Jin 0001, Yandan Wang, Cunjian Chen, Massimo Tistarelli
FG2
2021 "Who Wants to Know all this Stuff?!": Understanding Older Adults' Privacy Concerns in Aged Care Monitoring Devices
abstract
Abstract Aged care monitoring devices (ACMDs) enable older adults to live independently at home. But to do so, ACMDs collect and share older adults’ personal information with others, potentially raising privacy concerns. This paper presents a detailed account of the different privacy problems in ACMDs that concern older adults. We report findings from interviews and a focus group conducted with older adults who are ageing in place. Using Daniel Solove’s privacy taxonomy to categorize privacy concerns, our analysis suggests that older adults are concerned about the potential for ACMDs to give rise to six problems: surveillance, secondary use of data, breach of confidentiality, disclosure, decisional interference and disturbing others. Other findings indicate that participants are worried about their ability to impose control over collection and management of their personal details and are willing to only accept privacy trade-offs during emergencies. We provide recommendations for ACMD developers and future directions to address findings from this research.
Sami Alkhatib, Ryan Kelly 0001, Jenny Waycott, George Buchanan 0001, Marthie Grobler, Shuo Wang 0012
Interact. Comput.6
2020 PART-GAN: Privacy-Preserving Time-Series Sharing
Shuo Wang 0012, Carsten Rudolph, Surya Nepal, Marthie Grobler, Shangyu Chen
ICANN (1)1
2020 OIAD: One-for-all Image Anomaly Detection with Disentanglement Learning
abstract
Anomaly detection aims to recognize samples with anomalous and unusual patterns with respect to a set of normal data. This is significant for numerous domain applications, such as industrial inspection, medical imaging, and security enforcement. There are two key research challenges associated with existing anomaly detection approaches: (1) many approaches perform well on low-dimensional problems however the performance on high-dimensional instances, such as images, is limited; (2) many approaches often rely on traditional supervised approaches and manual engineering of features, while the topic has not been fully explored yet using modern deep learning approaches, even when the well-label samples are limited. In this paper, we propose a One-for-all Image Anomaly Detection system (OIAD) based on disentangled learning using only clean samples. Our key insight is that the impact of small perturbation on the latent representation can be bounded for normal samples while anomaly images are usually outside such bounded intervals, referred to as structure consistency. We implement this idea and evaluate its performance for anomaly detection. Our experiments with three datasets show that OIAD can detect over 90% of anomalies while maintaining a low false alarm rate. It can also detect suspicious samples from samples labeled as clean, coincided with what humans would deem unusual.
Shuo Wang 0012, Shangyu Chen, Surya Nepal, Carsten Rudolph, Marthie Grobler
IJCNN1
2020 Privacy-Preserving Data Generation and Sharing Using Identification Sanitizer
Shuo Wang 0012, Lingjuan Lyu, Shangyu Chen, Surya Nepal, Carsten Rudolph, Marthie Grobler
WISE (2)1
2019 Parametric Canonical Correlation Analysis
abstract
Generally, suppose a wave is a linear combination of multiple basis(Not necessarily a sine or cosine waves, it could also be a wavelet, etc.), different types of waves may be similar on some basis, but vary greatly on a certain basis. To address this problem, we introduce a PCCA-based feature extraction method that extends canonical correlation analysis (CCA). The PCCA-based method can train efficient classifiers to rely on only a few samples for periodic signals with support for removing noisy signals. As a demonstration, an efficient system is implemented for the classification of electrocardiogram (ECG) signals by PCCA. The performance is measured using several normal and abnormal ECG signals from the real-world database. These are compared with three commonly-adopted feature extraction techniques using five classes classification tasks related to ECG heartbeats. The AUC(Area under the ROC curve) of the PCCA-based feature extraction technique with two-digits size train dataset for four ECG type-pairs we compared were 0.8805, 0.957, 0.8968 and 1.00 respectively. The experimental results demonstrate that the proposed feature extraction techniques achieve better performance compared to other features extraction techniques with small amount of well-labeled data.
Shangyu Chen, Shuo Wang 0012, Richard O. Sinnott
CloudCom2
2019 P-STM: Privacy-Protected Social Tie Mining of Individual Trajectories
abstract
With the prevalence of location-aware devices and applications, enormous volumes of human spatiotemporal trajectories are being produced. It is feasible to estimate the similarity between user movement patterns according to such trajectories, which can be regarded as a potential social tie between users. There are two key research challenges associated with social tie discovery from trajectories: (1) trajectories contain users' accurate locations and releasing such data for social tie discovery raises serious privacy concerns; (2) trajectories are archived as discrete approximations of actual movement patterns using different sampling strategies and rates which are intrinsically heterogeneous. To address these challenges, this paper proposes a Privacy-protected Social Tie Mining (P-STM) approach. It provides a new social tie discovery solution based on the similarity of calibrated trajectories incorporating three key components: (1) a location entropy-based indicative dense region (IDR) mining approach to handle the heterogeneity of trajectories under differential privacy; (2) a private model-based calibration system used to rewrite trajectories using a sanitized IDR set to improve the utility of sanitized trajectories for similarity evaluation; (3) a social tie mining approach to indicate potential social ties between individuals using the similarity trajectories, which aims at finding the acquaintances for users based on solely their local geographical activities. The proposed approach is evaluated using real-world trajectory datasets from location-based social networks.
Shuo Wang 0012, Surya Nepal, Richard O. Sinnott, Carsten Rudolph
ICWS1
2018 Privacy-protected statistics publication over social media user trajectory streams
Shuo Wang 0012, Richard O. Sinnott, Surya Nepal
Future Gener. Comput. Syst.1
2017 Privacy-protected place of activity mining on big location data
abstract
People always spend their time at a few important locations for various activities in groups during specific time slots, called place of activity (POA), e.g., resting at home among family members during night and working at office among colleagues during work time. Inferring such places is significant for not only the precise advertising on the commercial aspect but the identifying rallies or meetings among a group of people and tracking of the target individuals on the aspect of public security, e.g., locating and tracking suspected terrorists for anti-terrorist work. However, it is a challenge to map from big location data to places of activity due to the volume and complexity whilst giving rise to privacy concerns, e.g., personally important place mining. In the paper, a method for POA mining on big location data is proposed, named P-PAM, aiming at big data analytics and privacy concerns. We use a clustering algorithm to discover the place of activity, then adopt location entropy as reference of user diversity and take into account temporal variation, to infer place of activity. Further, robust privacy-preserving mechanisms under differential privacy are embedded into clustering results and location entropy evaluation that accesses to raw location data. We demonstrate the utility of our proposed approach with large-scale location datasets derived from geo-referenced social media. The experimental results suggest that the POA mining approach can successfully scale to big data scenarios whilst preserving individual user privacy.
Shuo Wang 0012, Richard O. Sinnott, Surya Nepal
IEEE BigData1
2017 Sensitive gazetteer discovery and protection for mobile social media users
abstract
With the explosive growth of location-aware devices and global adoption of social network applications, enormous volumes of spatiotemporal data are being produced. These can be perceived as gazetteers that record frequently visited locations, e.g. shopping malls and museums, and potentially more sensitive locations, e.g. an individual's home/work locations. Density-based clustering approaches are generally used for gazetteer discovery. However, existing clustering solutions are inefficient for big data scenarios and often disregard mobility features derived from trajectories data. Further, automated gazetteer discovery applications may cause privacy concerns. In this paper, we propose a sensitive gazetteer automated discovery approach based on Ω-cluster with robust privacy controls. The approach identifies sensitive gazetteers from massive trajectory data, with location entropy-based filtering used to reduce the number of uninteresting clusters whilst considering mobility features of trajectories. A parallelized solution is implemented to scale across the cloud using memory-oriented data processing solutions based upon Apache Spark. We embed this algorithm in a privacy-preserving mechanism and subsequently release sanitized gazetteers. Through extensive experiments using synthetic and real trajectory datasets from the location based social network (Twitter), we demonstrate the effectiveness and efficiency of our approach.
Shuo Wang 0012, Richard O. Sinnott, Surya Nepal
IEEE BigData1
2017 Protecting personal trajectories of social media users through differential privacy
Shuo Wang 0012, Richard O. Sinnott
Comput. Secur.1
2016 Protecting the location privacy of mobile social media users
abstract
Unprecedented volumes of location-based information have been produced as a result of the widespread adoption of social network applications and GPS-enabled devices and sensors. Publication of such location data can provide valuable resources for researchers and government agencies in applications ranging from near real-time population-wide health monitoring to planning for future cities. However, such data hold personally identifying information, which gives rise to many privacy issues. There is thus a pressing need for ways to restrict this inherently identifying location-related information, however ideally we would like to preserve the utility of the data. Importantly, any such solution has to be scalable to large population-wide data scenarios. To tackle this, we introduce a novel differentially private hierarchical location sanitization (DPHLS) approach based on the concept “(α, r)-dataset” implemented through a Variable Order Mobility Markov Model (VO3M). We show how this system allows individual locations in personal trajectories to be protected using selection and frequency perturbation mechanisms using the “(α, r)-dataset”, leveraging past (published) location histories to obfuscate the user location in a flexible and controllable manner. The effectiveness and efficiency of the proposed solution is evaluated through the big data experiments that have been carried out using an OpenStack-based Cloud and Apache Sparkbased platform utilising large-scale social media trajectories. The experimental results suggest that the privacy publication algorithm can successfully scale to big data scenarios whilst retaining the utility of the datasets (trajectories) and preserving individual user privacy.
Shuo Wang 0012, Richard O. Sinnott, Surya Nepal
IEEE BigData1
2016 Privacy-protected social media user trajectories calibration
abstract
Advanced data analytics have become an integral part of a number of eScience initiatives including the many challenges facing the urban sciences. Understanding the movement of people and their spatial trajectorits would greatly aid the development of policies for sustainable urban living including urban traffic analysis and smart city management. Due to the widespread popularity of mobile devices with location-aware capabilities and the extensive prevalence of location-based social networks, citizens' movements and their trajectories are being produced and gathered at an unprecedented rate. In this context, there are two fundamental issues that need to be addressed: trajectory data hold private information of citizens that require privacy preserving solutions for data release and analysis, and the heterogeneous nature of the trajectory data makes it hard to effectively measure their similarity, which is fundamental to trajectory analysis. In this paper, we address these challenges by proposing an innovative private trajectories calibration model that not only guarantees the privacy of citizens, but also increases the utility. We have conducted comprehensive experiments using real-life user trajectories extracted from Twitter data. The results reveal the effectiveness and efficiency of the proposed approach, which is also reported in this paper.
Shuo Wang 0012, Richard O. Sinnott, Surya Nepal
eScience1
2006 Face-tracking as an augmented input in video games: enhancing presence, role-playing and control
abstract
Motion-detection only games have inherent limitations on game experience in that the systems cannot identify the player's existence and identity. A way of improvement is by introducing information such as a player's face or head into the system. We designed and implemented two game prototypes that apply real-time face position information as intrinsic elements of gameplay to enhance game experience. The first prototype augmented a typical motion-detection-based game. Face information was designed to enhance the sense of presence and role-playing. In the second prototype, face tracking is applied as a new axis of control in a First Person Shooter (FPS) game.Although Face detection and tracking technology has started utilizing in game scenarios, there was little systematic research on how user experience is leveraged by applying face information to video games. The results of our user tests on comparing camera-based video games with and without face tracking demonstrated that using face position information can effectively enhance presence and role-playing. In addition, an intuitive control that augmented by face-tracking in the FPS game also got positive feedbacks from the test.
Shuo Wang 0012, Xiaocao Xiong, Yan Xu 0011, Chao Wang 0063, Xiaofeng Dai, Dongmei Zhang 0001
CHI1