Hongsheng Hu

dblp:122/0826 · DBLP profile ↗
← Back
34ranked-venue papers
8as first author
32since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 3 first-author · 15 since 2021Security and privacy · 11 · 3 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 1 first-author · 10 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Systems, architecture and hardware · 2 · 2 first-author · 1 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Reference Recommendation Based Membership Inference Attack Against Hybrid-Based Recommender Systems
abstract
Recommender systems have been widely deployed across various domains such as e-commerce and social media, and intelligently suggest items like products and potential friends to users based on their preferences and interaction history, which are often privacy-sensitive. Recent studies have revealed that recommender systems are prone to membership inference attacks (MIAs), where an attacker aims to infer whether or not a user’s data has been used for training a target recommender system. However, existing MIAs fail to exploit the unique characteristic of recommender systems, and therefore are only applicable to mixed recommender systems consisting of two recommendation algorithms. This leaves a gap in investigating MIAs against hybrid-based recommender systems where the same algorithm utilizing user-item historical interactions and attributes of users and items serves and produces personalised recommendations. To investigate how the personalisation in hybrid-based recommender systems influences MIA, we propose a novel metric-based MIA. Specifically, we leverage the characteristic of personalisation to obtain reference recommendation for any target users. Then, a relative membership metric is proposed to exploit a target user’s historical interactions, target recommendation, and reference recommendation to infer the membership of the target user’s data. Finally, we theoretically and empirically demonstrate the efficacy of the proposed metric-based MIA on hybrid-based recommender systems.
Xiaoxiao Chi, Xuyun Zhang, Yan Wang 0002, Hongsheng Hu, Wan-Chun Dou
AAAI4
2026 ExpShield: Safeguarding Web Text from Unauthorized Crawling and LLM Exploitation
Ruixuan Liu, Tianhao Wang 0001, Hongsheng Hu, Shuo Wang 0012, Li Xiong 0001
NDSS4
2026 SoK: Robustness in Large Language Models against Jailbreak Attacks
Feiyue Xu, Hongsheng Hu, Chaoxiang He, Sheng Hang, Hanqing Hu, Zhengyan Zhou, Bin B. Zhu, Shifeng Sun 0001, Dawu Gu, Shuo Wang 0012
SP2
2026 Intellectual property protection for deep learning model and dataset intelligence
Yongqi Jiang, Yansong Gao 0001, Chunyi Zhou 0001, Hongsheng Hu, Anmin Fu, Willy Susilo
Eng. Appl. Artif. Intell.4
2026 From Pixels to Trajectory: Universal Adversarial Example Detection via Temporal Imprints
abstract
We unveil discernible temporal (or historical) trajectory imprints resulting from adversarial example (AE) attacks. Standing in contrast to existing studies, which focus on spatial (or static) imprints within the targeted underlying victim models, we present a novel temporal paradigm for understanding these attacks. These imprints are encapsulated within a single loss metric, spanning universally across diverse tasks such as classification and regression, and modalities including image, text, and audio. Recognizing the distinct nature of loss between adversarial and clean examples, we exploit this temporal imprint for AE detection by proposing (Traceable Adversarial Temporal Imprints). TRAIT operates under minimal assumptions without prior knowledge of attacks, thereby framing the detection challenge as a one-class classification problem. However, detecting AEs is still challenged by significant overlaps between the constructed synthetic losses of adversarial and clean examples due to the absence of ground truth for incoming inputs. TRAIT addresses this challenge by converting the synthetic loss into a spectrum signature, using the technique of Fast Fourier Transform to highlight the discrepancies, drawing inspiration from the temporal nature of the imprints, analogous to time-series signals. Across 12 AE attacks including SMACK (USENIX Sec'2023), TRAIT demonstrates consistent outstanding performance across comprehensively evaluated modalities (image, text, audio), tasks (classification and regression), datasets (nine datasets), and model architectures (e.g., ResNeXt50, BERT, RoBERTa, AudioNet). In all scenarios, TRAIT achieves an AE detection accuracy exceeding 97%, often around 99%, while maintaining a false rejection rate of 1%. TRAIT remains effective under the formulated strong adaptive attacks.
Yansong Gao 0001, Huaibing Peng, Zhiyang Dai, Shuo Wang 0012, Hongsheng Hu, Anmin Fu, Minhui Xue 0001
IEEE Trans. Dependable Secur. Comput.6
2025 When Better Features Mean Greater Risks: The Performance-Privacy Trade-Off in Contrastive Learning
abstract
When Better Features Mean Greater Risks: The Performance-Privacy Trade-Off in Contrastive Learning
Ruining Sun, Hongsheng Hu, Wei Luo 0001, Zhaoxi Zhang 0001, Yanjun Zhang 0002, Haizhuan Yuan, Leo Yu Zhang
AsiaCCS2
2025 Enhancing Adversarial Transferability with Checkpoints of a Single Model's Training
abstract
Adversarial attacks threaten the integrity of deep neural networks (DNNs), particularly in high-stakes applications. In this paper, we present a novel black-box adversarial attack that leverages the diverse checkpoints generated during a single model’s training trajectory. Unlike conventional ensemble attacks that require multiple surrogate models with diverse architectures, our approach exploits the intrinsic diversity captured over different training stages of a single surrogate model. By decomposing the learned representations into task-intrinsic and task-irrelevant components, we employ an accuracy gap-based selection strategy to identify checkpoints that predominantly capture transferable, task-intrinsic knowledge. Extensive experiments on ImageNet and CIFAR-10 demonstrate that our method consistently outperforms traditional ensemble attacks in terms of transferability, even under resource-constrained and practical settings. This work offers a resource-efficient solution for crafting highly transferable adversarial examples and provides new insights into the dynamics of adversarial vulnerability.
Shixin Li 0001, Chaoxiang He, Xiaojing Ma 0002, Bin B. Zhu, Shuo Wang 0012, Hongsheng Hu, Dongmei Zhang 0001, Linchen Yu
CVPR6
2025 Where Does This Data Come From? Enhanced Source Inference Attacks in Federated Learning
abstract
Federated learning (FL) enables collaborative model training without exposing raw data, offering a privacy-aware alternative to centralized learning. However, FL remains vulnerable to various privacy attacks that exploit shared model updates, including membership inference, property inference, and gradient inversion. Source inference attacks further threaten FL by identifying which client contributed a specific training sample, posing severe risks to user and institutional privacy. Existing source inference attacks mainly assume passive adversaries and overlook more realistic scenarios where the server actively manipulates the training process. In this paper, we present an enhanced source inference attack that demonstrates how a malicious server can amplify behavioral differences between clients to more accurately infer data origin. Our approach introduces active training manipulation and data augmentation to expose client-specific patterns. Experimental results across five representative FL algorithms and multiple datasets show that our method significantly outperforms prior passive attacks. These findings reveal a deeper level of privacy vulnerability in FL and call for stronger defense mechanisms under active threat models.
Xiaolong Xu 0001, Xiaokang Zhou, Fei Dai 0002, Yansong Gao 0001, Shuo Wang 0026, Hongsheng Hu
IJCAI9
2025 Fine-Grained and Efficient Self-Unlearning with Layered Iteration
abstract
As machine learning models become widely deployed in data-driven applications, ensuring compliance with the 'right to be forgotten' as required by many privacy regulations is vital for safeguarding user privacy. To forget the given data, existing re-labeling based unlearning methods employ a single-step adjustment scheme that revises the decision boundaries in one re-labeling phase. However, such single-step approaches lead to coarse-grained changes in decision boundaries among the remaining classes and impose adverse effects on the model utility. To address these limitations, we propose 'Self-Unlearning with Layered Iteration (SULI),' a novel unlearning approach that introduces a layered iteration strategy to re-label the forgetting data iteratively and refine the decision boundaries progressively. We further develop a 'Selective Probability Adjustment (SPA)' technique, which uses a soft-label mechanism to promote smoother decision-boundary transitions. Comprehensive experiments on three benchmark datasets demonstrate that SULI achieves superior performance in effectiveness, efficiency, and privacy compared to the state-of-the-art baselines in both class-wise and instance-wise unlearning scenarios. The source code is released at https://github.com/Hongyi-Lyu-MQ/SULI.
Hongyi Lyu, Xuyun Zhang, Hongsheng Hu, Shuo Wang 0026, Chaoxiang He, Lianyong Qi
IJCAI3
2025 Balancing Invariant and Specific Knowledge for Domain Generalization with Online Knowledge Distillation
abstract
Recent research has demonstrated the effectiveness of knowledge distillation in Domain Generalization. However, existing approaches often overlook domain-specific knowledge and rely on an offline distillation strategy, limiting the effectiveness of knowledge transfer. To address these limitations, we propose Balanced Online knowLedge Distillation (BOLD). BOLD leverages a multi-domain expert teacher model, with each expert specializing in a specific source domain, enabling the student to distill both domain-invariant and domain-specific knowledge. We incorporate the Pareto optimization principle and uncertainty weighting to balance these two types of knowledge, ensuring simultaneous optimization without compromising either. Additionally, BOLD employs an online knowledge distillation strategy, allowing the teacher and student to learn concurrently. This dynamic interaction enables the teacher to adapt based on student feedback, facilitating more effective knowledge transfer. Extensive experiments on seven benchmarks demonstrate that BOLD outperforms state-of-the-art methods. Furthermore, we provide theoretical insights that highlight the importance of domain-specific knowledge and the advantages of uncertainty weighting.
Jingfeng Zhang, Hongsheng Hu, Philippe Fournier-Viger, Gillian Dobbie, Yun Sing Koh
IJCAI3
2025 Boosting Guided Diffusion with Large Language Models for Multimodal Sequential Recommendation
abstract
Recent advancements in generative models have positioned them as one of the principal tools for sequential recommendation due to their exceptional sample diversity and generalization capabilities. Among these, diffusion model-based sequential recommenders have achieved remarkable success. However, most existing approaches still face critical challenges, resulting in suboptimal generation quality: (1) They fail to leverage multimodal knowledge for constructing item representations with well-structured distributional characteristics and semantically enriched information; (2) They predominantly rely on discrete diffusion processes, leading to high error accumulation, reduced time efficiency, and constrained controllability in generative sampling. To mitigate these challenges, we propose LSGM4Rec, a novel framework that integrates Large Language Models (LLMs) with advanced multimodal encoding models to establish multimodal fusion embeddings for items. This design ensures distinct distributional characteristics while enabling the incorporation of semantically rich modal features into guidance condition. Furthermore, we pioneer the stochastic differential equations (SDEs) for recommendation, facilitating smooth transitions between data distributions and enabling optimal trade-off between sampling efficiency and generation quality. Extensive experiments on three datasets demonstrate that LSGM4Rec outperforms existing state-of-the-art sequential recommendation methods.
Te Song, Lianyong Qi, Weiming Liu 0005, Fan Wang 0020, Xiaolong Xu 0001, Hongsheng Hu, Yang Cao 0019, Xuyun Zhang, Amin Beheshti
ACM Multimedia6
2025 HPSERec: A Hierarchical Partitioning and Stepwise Enhancement Framework for Long-tailed Sequential Recommendation
abstract
The long-tail problem in sequential recommender systems stems from imbalanced interaction data, resulting in suboptimal model performance for tail users and items. Recent studies have leveraged head data to enhance tail data for diminish the impact of the long-tail problem. However, these methods often adopt ad-hoc strategies to distinguish between head and tail data, which fails to capture the underlying distributional characteristics and structural properties of each category. Moreover, due to a substantial representational gap exists between head and tail data, head-to-tail enhancement strategies are susceptible to negative transfer, often leading to a decline in overall model performance. To address these issues, we propose a hierarchical partitioning and stepwise enhancement framework, called HPSERec, for long-tailed sequential recommendation. HPSERec partitions the item set into subsets based on a data imbalance metric, assigning an expert network to each subset to capture user-specific local features. Subsequently, we apply knowledge distillation to progressively improve long-tail interest representation, followed by a Sinkhorn optimal transport-based feedback module, which aligns user representations across expert levels through a globally optimal and softly matched mapping. Extensive experiments on three real-world datasets demonstrate that HPSERec consistently outperforms all baseline methods. The implementation code is available at https://anonymous.4open.science/r/HPSERec-2404.
Xiaolong Xu 0001, Haolong Xiang, Xuyun Zhang, Wei Shen 0005, Hongsheng Hu, Lianyong Qi
NeurIPS6
2025 BadFU: Backdoor Federated Learning through Adversarial Machine Unlearning
abstract
Federated learning (FL) has been widely adopted as a decentralized training paradigm that enables multiple clients to collaboratively learn a shared model without exposing their local data. As concerns over data privacy and regulatory compliance grow, machine unlearning, which aims to remove the influence of specific data from trained models, has become increasingly important in the federated setting to meet legal, ethical, or user-driven demands. However, integrating unlearning into FL introduces new challenges and raises largely unexplored security risks. In particular, adversaries may exploit the unlearning process to compromise the integrity of the global model. In this paper, we present the first backdoor attack in the context of federated unlearning, demonstrating that an adversary can inject backdoors into the global model through seemingly legitimate unlearning requests. Specifically, we propose BadFU, an attack strategy where a malicious client uses both backdoor and camouflage samples to train the global model normally during the federated training process. Once the client requests unlearning of the camouflage samples, the global model transitions into a backdoored state. Extensive experiments under various FL frameworks and unlearning strategies validate the effectiveness of BadFU, revealing a critical vulnerability in current federated unlearning practices and underscoring the urgent need for more secure and robust federated unlearning mechanisms.
Bingguang Lu, Hongsheng Hu, Yuantian Miao, Shaleeza Sohail, Chaoxiang He, Shuo Wang 0012, Xiao Chen 0002
RAID2
2025 MGF-ESE: An Enhanced Semantic Extractor with Multi-Granularity Feature Fusion for Code Summarization
abstract
Code summarization aims to generate concise natural language descriptions of source code, helping developers to acquaint with software systems and reduce maintenance costs. Existing code summarization approaches widely employ attention mechanisms to assess the relevance between nodes in the Abstract Syntax Tree (AST), which generates context vectors that reflect the semantics of the source code. However, these approaches solely relying on AST lack the extraction of features at other levels of granularity, such as code tokens and Control Flow Graph (CFG), which suffer from severe semantic gaps when capturing data and control dependencies. To address this issue, we design an enhanced semantic extractor with multi-granularity feature fusion (MGF-ESE) to improve the model capability in comprehending and processing the overall semantics of the code. Specifically, we present a novel AST generation method that, based on controlling the scale of nodes, introduces syntactic description nodes to raise the semantic density of AST feature. Then we perform both local and global encoding of CFG after embedding the statement nodes. Moreover, through a cross-attention mechanism, we fuse code tokens and CFG with AST to enhance the model's capacity to capture both syntactic and structural information from source code. Finally, extensive experiments on two open-source datasets show that MGF-ESE outperforms the state-of-the-arts with higher-quality code summaries on key metrics, including BLEU, METEOR, and ROUGE-L.
Xiaolong Xu 0001, Hongsheng Hu, Haolong Xiang, Lianyong Qi, Junqun Xiong, Wan-Chun Dou
WWW3
2025 Artificial intelligence security and privacy: a survey
abstract
Abstract Artificial intelligence (AI) is revolutionizing both industries and reshaping the global economy. However, the rapid advancement of AI technologies brings significant security and privacy challenges. Recent incidents highlight vulnerabilities in AI systems, such as data leakage and malicious code injection, leading to severe financial losses and privacy breaches. Although existing studies have discussed specific security threats, they often lack detailed granularity and cover a limited scope. In this survey, we fill this gap by systematically categorizing and analyzing the threats and countermeasures in AI systems, which span both the training and inference stages, encompass centralized and distributed settings, and address both conventional and foundation AI models. By reviewing existing literature, we aim to provide AI researchers and practitioners with a thorough understanding of system vulnerabilities and current countermeasures. We hope to inspire further research into robust solutions, ultimately contributing to the development of resilient AI technologies.
Xinlei He 0001, Guowen Xu, Xingshuo Han, Qian Wang 0002, Lingchen Zhao, Chao Shen 0001, Chenhao Lin, Zhengyu Zhao 0001, Qian Li 0024, Le Yang 0007, Shouling Ji, Shaofeng Li 0001, Haojin Zhu, Zhibo Wang 0001, Tianqing Zhu, Qi Li 0002, Chaoxiang He, Hongsheng Hu, Shuo Wang 0012, Shifeng Sun 0001, Hongwei Yao, Qinyu Zhang 0001, Kai Chen 0012, Yue Zhao 0027, Hongwei Li 0001, Xinyi Huang 0001, Dengguo Feng
Sci. China Inf. Sci.20
2024 Symmetric Self-Paced Learning for Domain Generalization
abstract
Deep learning methods often suffer performance degradation due to domain shift, where discrepancies exist between training and testing data distributions. Domain generalization mitigates this problem by leveraging information from multiple source domains to enhance model generalization capabilities for unseen domains. However, existing domain generalization methods typically present examples to the model in a random manner, overlooking the potential benefits of structured data presentation. To bridge this gap, we propose a novel learning strategy, Symmetric Self-Paced Learning (SSPL), for domain generalization. SSPL consists of a Symmetric Self-Paced training scheduler and a Gradient-based Difficulty Measure (GDM). Specifically, the proposed training scheduler initially focuses on easy examples, gradually shifting emphasis to harder examples as training progresses. GDM dynamically evaluates example difficulty through the gradient magnitude with respect to the example itself. Experiments across five popular benchmark datasets demonstrate the effectiveness of the proposed learning strategy.
Yun Sing Koh, Gillian Dobbie, Hongsheng Hu, Philippe Fournier-Viger
AAAI4
2024 Shadow-Free Membership Inference Attacks: Recommender Systems Are More Vulnerable Than You Thought
Xiaoxiao Chi, Xuyun Zhang, Yan Wang 0002, Lianyong Qi, Amin Beheshti, Xiaolong Xu 0001, Kim-Kwang Raymond Choo, Shuo Wang 0026, Hongsheng Hu
IJCAI9
2024 A Duty to Forget, a Right to be Assured? Exposing Vulnerabilities in Machine Unlearning Services
Hongsheng Hu, Shuo Wang 0012, Jiamin Chang, Haonan Zhong, Ruoxi Sun 0001, Shuang Hao 0001, Haojin Zhu, Minhui Xue 0001
NDSS1
2024 DapperFL: Domain Adaptive Federated Learning with Model Fusion Pruning for Edge Devices
abstract
Federated learning (FL) has emerged as a prominent machine learning paradigm in edge computing environments, enabling edge devices to collaboratively optimize a global model without sharing their private data. However, existing FL frameworks suffer from efficacy deterioration due to the system heterogeneity inherent in edge computing, especially in the presence of domain shifts across local data. In this paper, we propose a heterogeneous FL framework DapperFL, to enhance model performance across multiple domains. In DapperFL, we introduce a dedicated Model Fusion Pruning (MFP) module to produce personalized compact local models for clients to address the system heterogeneity challenges. The MFP module prunes local models with fused knowledge obtained from both local and remaining domains, ensuring robustness to domain shifts. Additionally, we design a Domain Adaptive Regularization (DAR) module to further improve the overall performance of DapperFL. The DAR module employs regularization generated by the pruned model, aiming to learn robust representations across domains. Furthermore, we introduce a specific aggregation algorithm for aggregating heterogeneous local models with tailored architectures and weights. We implement DapperFL on a real-world FL platform with heterogeneous clients. Experimental results on benchmark datasets with multiple domains demonstrate that DapperFL outperforms several state-of-the-art FL frameworks by up to 2.28%, while significantly achieving model volume reductions ranging from 20% to 80%. Our code is available at: https://github.com/jyzgh/DapperFL.
Yongzhe Jia, Xuyun Zhang, Hongsheng Hu, Kim-Kwang Raymond Choo, Lianyong Qi, Xiaolong Xu 0001, Amin Beheshti, Wan-Chun Dou
NeurIPS3
2024 Learn What You Want to Unlearn: Unlearning Inversion Attacks against Machine Unlearning
abstract
Machine unlearning has become a promising solution for fulfilling the "right to be forgotten", under which individuals can request the deletion of their data from machine learning models. However, existing studies of machine unlearning mainly focus on the efficacy and efficiency of unlearning methods, while neglecting the investigation of the privacy vulnerability during the unlearning process. With two versions of a model available to an adversary, that is, the original model and the unlearned model, machine unlearning opens up a new attack surface. In this paper, we conduct the first investigation to understand the extent to which machine unlearning can leak the confidential content of the unlearned data. Specifically, under the Machine Learning as a Service setting, we propose unlearning inversion attacks that can reveal the feature and label information of an unlearned sample by only accessing the original and unlearned model. The effectiveness of the proposed unlearning inversion attacks is evaluated through extensive experiments on benchmark datasets across various model architectures and on both exact and approximate representative unlearning approaches. The experimental results indicate that the proposed attack can reveal the sensitive information of the unlearned data. As such, we identify three possible defenses that help to mitigate the proposed attacks, while at the cost of reducing the utility of the unlearned model. The study in this paper uncovers an underexplored gap between machine unlearning and the privacy of unlearned data, highlighting the need for the careful design of mechanisms for implementing unlearning without leaking the information of the unlearned data.
Hongsheng Hu, Shuo Wang 0012, Tian Dong 0003, Minhui Xue 0001
SP1
2024 LACMUS: Latent Concept Masking for General Robustness Enhancement of DNNs
abstract
The susceptibility of Deep Neural Networks (DNNs) to adversarial attacks and their limited robustness to real-world variations pose substantial challenges to their widespread adoption. Adversarial training has shown promise in fortifying models against such perturbations, however current methods are often specific to a single type of attack and can significantly diminish the model’s overall performance. In response, we present LAtent Concept Masking for robUStness (LACMUS), a novel perceptually-driven methodology that enhances DNN robustness without requiring prior knowledge about the adversarial contexts. We argue that DNNs’ sensitivity to adversarial perturbations and distribution drifts stems from overfitting to non-common concepts within the dataset, leading to an over-reliance on specific learned instances and increased vulnerability. LACMUS addresses this by mapping high-dimensional data into a latent conceptual space to identify and navigate patterns of "non-common concepts" within the latent concept space. It then applies a concept masking strategy to selectively obscure data features, prompting the model to base its decisions on a wider array of information and thus enhancing its decision-making robustness. LACMUS distinguishes itself as a versatile, attack-agnostic framework that employs concept-wise augmentation to enhance robustness against a spectrum of adversarial, semantic, and distributional challenges. Our contributions include the development of a tool for robustness enhancement, a mechanism for mapping data to latent concept space, a strategy for identifying patterns of concept-wise misclassification, and a novel data augmentation module that leverages latent concepts. LACMUS is proven to enhance model resilience and generalization, even when training data is scarce, with experiments on MNIST, CIFAR-10, ImageNet, and CelebA supporting its effectiveness. We also provide augmented datasets to the research community, bolstering the robustness of models trained on them.
Shuo Wang 0012, Hongsheng Hu, Jiamin Chang, Benjamin Zi Hao Zhao, Minhui Xue 0001
SP2
2024 DNN-GP: Diagnosing and Mitigating Model's Faults Using Latent Concepts
Shuo Wang 0012, Hongsheng Hu, Jiamin Chang, Benjamin Zi Hao Zhao, Qi Alfred Chen, Minhui Xue 0001
USENIX Security Symposium2
2024 Cardinality Counting in "Alcatraz": A Privacy-aware Federated Learning Approach
abstract
The task of cardinality counting, pivotal for data analysis, endeavors to quantify unique elements within datasets and has significant applications across various sectors like healthcare, marketing, cybersecurity, and web analytics. Current methods, categorized into deterministic and probabilistic, often fail to prioritize data privacy. Given the fragmentation of datasets across various organizations, there is an elevated risk of inadvertently disclosing sensitive information during collaborative data studies using state-of-the-art cardinality counting techniques. This study introduces an innovative privacy-centric solution for the cardinality counting dilemma, leveraging a federated learning framework. Our approach involves employing a locally differentially private data encoding for initial processing, followed by a privacy-aware federated K-means clustering strategy, ensuring that cardinality counting occurs across distinct datasets without necessitating data amalgamation. The efficacy of our methodology is underscored by promising results from tests on both real-world and simulated datasets, pointing towards a transformative approach to privacy-sensitive cardinality counting in contemporary data science.
Nan Wu 0013, Xin Yuan 0004, Shuo Wang 0012, Hongsheng Hu, Minhui Xue 0001
WWW4
2024 Source Inference Attacks: Beyond Membership Inference Attacks in Federated Learning
abstract
Federated learning (FL) is a popular approach to facilitate privacy-aware machine learning since it allows multiple clients to collaboratively train a global model without granting others access to their private data. It is, however, known that FL can be vulnerable tomembership inference attacks(MIAs), where the training records of the global model can be distinguished from the testing records. Surprisingly, research focusing on the investigation of the source inference problem appears to be lacking. We also observe that identifying a training record's source client can result in privacy breaches extending beyond MIAs. For example, consider an FL application where multiple hospitals jointly train a COVID-19 diagnosis model, membership inference attackers can identify the medical records that have been used for training, and any additional identification of the source hospital can result the patient from the particular hospital more prone to discrimination. Seeking to contribute to the literature gap, we take the first step to investigate source privacy in FL. Specifically, we propose a new inference attack (hereafter referred to assource inference attack– SIA), designed to facilitate an honest-but-curious server to identify the training record's source client. The proposed SIAs leverage the Bayesian theorem to allow the server to implement the attack in a non-intrusive manner without deviating from the defined FL protocol. We then evaluate SIAs in three different FL frameworks to show that in existing FL frameworks, the clients sharing gradients, model parameters, or predictions on a public dataset will leak such source information to the server. We also conduct extensive experiments on various datasets to investigate the key factors in an SIA. The experimental results validate the efficacy of the proposed SIAs, e.g., an attack success rate of 67.1% (baseline 10%) can be achieved when the clients share model parameters with the server. Comprehensive ablation studies demonstrate that the success of an SIA is directly related to the overfitting of the local models.
Hongsheng Hu, Xuyun Zhang, Zoran A. Salcic, Lichao Sun 0001, Kim-Kwang Raymond Choo, Gillian Dobbie
IEEE Trans. Dependable Secur. Comput.1
2023 POSTER: ML-Compass: A Comprehensive Assessment Framework for Machine Learning Models
abstract
Machine learning models have made significant breakthroughs across various domains. However, it is crucial to assess these models to obtain a complete understanding of their capabilities and limitations and ensure their effectiveness and reliability in solving real-world problems. In this paper, we present a framework, termed ML-Compass, that covers a broad range of machine learning abilities, including utility evaluation, neuron analysis, robustness evaluation, and interpretability examination. We use this framework to assess seven state-of-the-art classification models on four benchmark image datasets. Our results indicate that different models exhibit significant variation, even when trained on the same dataset. This highlights the importance of using the assessment framework to comprehend their behavior.
Zhibo Jin, Hongsheng Hu, Minhui Xue 0001, Huaming Chen
AsiaCCS3
2023 OptIForest: Optimal Isolation Forest for Anomaly Detection
abstract
Anomaly detection plays an increasingly important role in various fields for critical tasks such as intrusion detection in cybersecurity, financial risk detection, and human health monitoring. A variety of anomaly detection methods have been proposed, and a category based on the isolation forest mechanism stands out due to its simplicity, effectiveness, and efficiency, e.g., iForest is often employed as a state-of-the-art detector for real deployment. While the majority of isolation forests use the binary structure, a framework LSHiForest has demonstrated that the multi-fork isolation tree structure can lead to better detection performance. However, there is no theoretical work answering the fundamentally and practically important question on the optimal tree structure for an isolation forest with respect to the branching factor. In this paper, we establish a theory on isolation efficiency to answer the question and determine the optimal branching factor for an isolation tree. Based on the theoretical underpinning, we design a practical optimal isolation forest OptIForest incorporating clustering based learning to hash which enables more information to be learned from data for better isolation quality. The rationale of our approach relies on a better bias-variance trade-off achieved by bias reduction in OptIForest. Extensive experiments on a series of benchmarking datasets for comparative and ablation studies demonstrate that our approach can efficiently and robustly achieve better detection performance in general than the state-of-the-arts including the deep learning based methods.
Haolong Xiang, Xuyun Zhang, Hongsheng Hu, Lianyong Qi, Wan-Chun Dou, Mark Dras, Amin Beheshti, Xiaolong Xu 0001
IJCAI3
2023 Differentially private locality sensitive hashing based federated recommender system
abstract
Summary Recommender systems are important applications in big data analytics because accurate recommendation items or high‐valued suggestions can bring high profit to both commercial companies and customers. To make precise recommendations, a recommender system often needs large and fine‐grained data for training. In the current big data era, data often exists in the form of isolated islands, and it is difficult to integrate the data scattered due to privacy security concerns. Moreover, privacy laws and regulations make it harder to share data. Therefore, designing a privacy‐preserving recommender system is of paramount importance. Existing privacy‐preserving recommender system models mainly adapt cryptography approaches to achieve privacy preservation. However, cryptography approaches have heavy overhead when performing encryption and decryption operations and they lack a good level of flexibility. In this paper, we conduct privacy analysis on the existing locality sensitive hashing (LSH) approach based privacy‐preserving recommender system and show how an attacker can retrieve user's information under such a recommender system. Given such privacy risks, we propose differentially private LSH approach to build recommender system that can offer differential privacy guarantees for users. Our proposed efficient and scalable federated recommender system can make full use of multiple source data from different data owners while guaranteeing privacy preservation of users' data in contributing parties. Extensive experiments on real‐world benchmark datasets show that our approach can achieve both high time efficiency and accuracy under small privacy budgets.
Hongsheng Hu, Gillian Dobbie, Zoran A. Salcic, Meng Liu 0007, Lingjuan Lyu, Xuyun Zhang
Concurr. Comput. Pract. Exp.1
2022 DeepiForest: A Deep Anomaly Detection Framework with Hashing Based Isolation Forest
abstract
With the great success of deep neural networks (DNNs) in a variety of fields, deep learning gains a pioneering development in anomaly detection. Although deep learning achieves good accuracy in anomaly detection, it is troubled with long execution time and high memory consumption. These problems are associated with the inherent drawbacks of deep learning, such as too many parameters and deep training layers. To remedy the above drawbacks, we try to explore an unsupervised non-neural network deep model for anomaly detection based on the experience of the deep forest. In this paper, we propose a deep anomaly detection framework with hashing based isolation forest (DeepiForest) to achieve effective and robust anomaly detection. Specifically, DeepiForest utilizes hashing based isolation forest and tree-embedding scheme to provide enhanced features and apply multi-layer cascaded architecture to establish a deep framework. DeepiForest inherits the advantages of deep forests, i.e., the framework holds fewer hyper-parameters and smaller model complexity than DNNs, simultaneously producing robust accuracy on anomaly detection. Extensive experiments on different-scale datasets illustrate the efficiency of DeepiForest and its comparable effectiveness to the state-of-the-art deep anomaly detection (DAD) methods.
Haolong Xiang, Hongsheng Hu, Xuyun Zhang
ICDM2
2022 Membership Inference via Backdooring
abstract
Recently issued data privacy regulations like GDPR (General Data Protection Regulation) grant individuals the right to be forgotten. In the context of machine learning, this requires a model to forget about a training data sample if requested by the data owner (i.e., machine unlearning). As an essential step prior to machine unlearning, it is still a challenge for a data owner to tell whether or not her data have been used by an unauthorized party to train a machine learning model. Membership inference is a recently emerging technique to identify whether a data sample was used to train a target model, and seems to be a promising solution to this challenge. However, straightforward adoption of existing membership inference approaches fails to address the challenge effectively due to being originally designed for attacking membership privacy and suffering from several severe limitations such as low inference accuracy on well-generalized models. In this paper, we propose a novel membership inference approach inspired by the backdoor technology to address the said challenge. Specifically, our approach of Membership Inference via Backdooring (MIB) leverages the key observation that a backdoored model behaves very differently from a clean model when predicting on deliberately marked samples created by a data owner. Appealingly, MIB requires data owners' marking a small number of samples for membership inference and only black-box access to the target model, with theoretical guarantees for inference results. We perform extensive experiments on various datasets and deep neural network architectures, and the results validate the efficacy of our approach, e.g., marking only 0.1% of the training dataset is practically sufficient for effective membership inference.
Hongsheng Hu, Zoran A. Salcic, Gillian Dobbie, Jinjun Chen, Lichao Sun 0001, Xuyun Zhang
IJCAI1
2021 Source Inference Attacks in Federated Learning
abstract
Federated learning (FL) has emerged as a promising privacy-aware paradigm that allows multiple clients to jointly train a model without sharing their private data. Recently, many studies have shown that FL is vulnerable to membership inference attacks (MIAs) that can distinguish the training members of the given model from the non-members. However, existing MIAs ignore the source of a training member, i.e., the information of the client owning the training member, while it is essential to explore source privacy in FL beyond membership privacy of examples from all clients. The leakage of source information can lead to severe privacy issues. For example, identification of the hospital contributing to the training of an FL model for the COVID-19 pandemic can render the owner of a data record from this hospital more prone to discrimination if the hospital is in a high risk region. In this paper, we propose a new inference attack called source inference attack (SIA), which can derive an optimal estimation of the source of a training member. Specifically, we innovatively adopt the Bayesian perspective to demonstrate that an honest-but-curious server can launch an SIA to steal non-trivial source information of the training members without violating the FL protocol. The server leverages the prediction loss of local models on the training members to achieve the attack effectively and non-intrusively. We conduct extensive experiments on one synthetic and five real datasets to evaluate the key factors in an SIA, and the results show the efficacy of the proposed source inference attack.
Hongsheng Hu, Zoran A. Salcic, Lichao Sun 0001, Gillian Dobbie, Xuyun Zhang
ICDM1
2021 EAR: An Enhanced Adversarial Regularization Approach against Membership Inference Attacks
abstract
Membership inference attacks on a machine learning model aim to determine whether a given data record is a member of the training set. They pose severe privacy risks to individuals, e.g., identifying an individual's participation in a hospital's health analytic training set reveals that this individual was once a patient in that hospital. Adversarial regularization (AR) is one of the state-of-the-art defense methods that mitigate such attacks while preserving a model's prediction accuracy. AR adds membership inference attacks as a new regularization term to the target model during the training process. It is an adversarial training algorithm that is trained on a defended model which is essentially the same as training the generator of generative adversarial networks (GANs). We observe that many GAN variants are able to generate higher quality samples and offer more stability during the training phase than GANs. However, whether these GAN variants are available to improve the effectiveness of AR has not been investigated. In this paper, we propose an enhanced adversarial regularization (EAR) method based on Least Square GANs (LSGANs). The new EAR surpasses the existing AR in offering more powerful defensive ability while preserving the same prediction accuracy of the protected classifiers. We systematically evaluate EAR on five datasets with different target classifiers under four different attack methods and compare it with four other defense methods. We experimentally show that our new method performs the best among other defense methods.
Hongsheng Hu, Zoran A. Salcic, Gillian Dobbie, Yi Chen 0008, Xuyun Zhang
IJCNN1
2021 Clustering-based Efficient Privacy-preserving Face Recognition Scheme without Compromising Accuracy
abstract
Recently, biometric identification has been extensively used for border control. Some face recognition systems have been designed based on Internet of Things. But the rich personal information contained in face images can cause severe privacy breach and abuse issues during the process of identification if a biometric system has compromised by insiders or external security attacks. Encrypting the query face image is the state-of-the-art solution to protect an individual’s privacy but incurs huge computational cost and poses a big challenge on time-critical identification applications. However, due to their high computational complexity, existing methods fail to handle large-scale biometric repositories where a target face is searched. In this article, we propose an efficient privacy-preserving face recognition scheme based on clustering. Concretely, our approach innovatively matches an encrypted face query against clustered faces in the repository to save computational cost while guaranteeing identification accuracy via a novel multi-matching scheme. To the best of our knowledge, our scheme is the first to reduce the computational complexity from O(M) in existing methods to approximate O (√ M ), where M is the size of a face repository. Extensive experiments on real-world datasets have shown the effectiveness and efficiency of our scheme.
Meng Liu 0007, Hongsheng Hu, Haolong Xiang, Chi Yang, Lingjuan Lyu, Xuyun Zhang
ACM Trans. Sens. Networks2
2020 A Locality Sensitive Hashing Based Approach for Federated Recommender System
abstract
The recommender system is an important application in big data analytics because accurate recommendation items or high-valued suggestions can bring high profit to both commercial companies and customers. To make precise recommendations, a recommender system often needs large and fine-grained data for training. In the current big data era, data often exist in the form of isolated islands, and it is difficult to integrate the data scattered due to privacy security concerns. Moreover, privacy laws and regulations make it harder to share data. Therefore, designing a privacy-preserving recommender system is of paramount importance. Existing privacy-preserving recommender system models mainly adapt cryptography approaches to achieve privacy preservation. However, cryptography approaches have heavy overhead when performing encryption and decryption operations and they lack a good level of flexibility. In this paper, we propose a Locality Sensitive Hashing (LSH) based approach for federated recommender system. Our proposed efficient and scalable federated recommender system can make full use of multiple source data from different data owners while guaranteeing preservation of privacy of contributing parties. Extensive experiments on real-world benchmark datasets show that our approach can achieve both high time efficiency and accuracy under small privacy budgets.
Hongsheng Hu, Gillian Dobbie, Zoran A. Salcic, Meng Liu 0007, Xuyun Zhang
CCGRID1
2008 Research on water quality assessment method based on multi-class support vector machines
abstract
It has been a more complex problem for water quality assessment. And its aim is to well and truly evaluate its degree of pollution for bodies of water, which will be easy to provide some principled projects and criterions for water resource's protection and their integration application. So, it has been widely applied into water quality assessment. SVM and directed acyclic graph support vector machine (DAGSVM) are paid much attention in this paper. A water quality assessment method based on DAGSVM is put forward. The test results show that the method proposed in this paper has an excellent performance on correct ratio. Besides, a wireless water quality monitoring system is designed and developed. The system can realize the acquisition of the water parameters, and the acquisition information timely and low cost transmission by GPRS (general packet radio service, GPRS). The developed system can store and display the water quality parameters from the monitoring spots in real time. Combining the wireless communication technology with the monitoring technology, the designed and developed system can greatly improve the real-time and continuity for the water quality's monitoring and assessment.
Hongsheng Hu, Suxiang Qian, Xiaojun Gu
ICARCV2