Siquan Huang

dblp:332/1734 · DBLP profile ↗
← Back
11ranked-venue papers
5as first author
11since 2021 · last 2026
0000-0003-0648-3405ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021Security and privacy · 4 · 1 first-author · 4 since 2021Computer networks · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 A website fingerprinting attack with unsupervised out-of-distribution detection
abstract
Abstract Website fingerprinting attacks are critical for extracting website information and identifying illegal websites visited by users in anonymous networks such as Tor. However, existing attacks struggle to extract effective features from unmonitored websites due to the diversity. Although increasing unmonitored training data can improve effectiveness, it also increases attacker costs. To address this, we propose a novel website fingerprinting attack that leverages unsupervised Out-of-Distribution detection. We exclusively use monitored website data for model training, eliminating the need for extensive unmonitored samples. For feature extraction, we utilize a combination of Long Short-Term Memory and Convolutional Neural Networks for robust feature extraction of each monitored website. We also introduce a new loss function to maximize differentiation between features of various websites. Furthermore, we employ Singular Value Decomposition to effectively segregate monitored from unmonitored websites. It allows the model to focus on dominant components in the feature vectors, facilitating a clear distinction between monitored and unmonitored website traffic. The experimental results confirm that our method outperforms existing techniques without requiring unmonitored training data.
Ying Gao 0004, Jiafeng Zhao, Chong Chen 0011, Siquan Huang, Leyu Shi, Chenglong Jiang
Cybersecur.4
2026 VulSCC: image-based vulnerability detection with SPP-CNN and code large language model
abstract
Abstract Deep learning excels in detecting source code vulnerabilities, where image-based detection methods overcome ignoring deep code semantic information in token-based methods and the inefficiency of graph-based methods. Unfortunately, current image-based methods cannot sufficiently extract vulnerability-related features due to three key limitations: (1) the inappropriateness of the construction of node centrality for sequential Program Dependency Graphs (PDGs), (2) the ineffective code analysis of traditional embedding models, and (3) the poor existing truncation/padding methods. Moreover, they fail to achieve effective vulnerability localization due to the irregular output of their interpretation method. In response, we propose a novel image-based line-level source code vulnerability detection system VulSCC. Firstly, VulSCC constructs a novel centrality combination and leverages a code large language model to capture richer vulnerability-related features. Secondly, we integrate an SPP layer to convert PDGs into images without distortion, as it adaptively aggregates arbitrary sizes into fixed-length vectors without truncation/padding. Finally, we use the occlusion technique to interpret the model predictions, which locate specific vulnerability lines, enabling effective vulnerability localization. Experimental results of VulSCC against seven SoTA methods show optimal detection performance in function-level detection. Additionally, we evaluate the effectiveness of the occlusion technique in localizing vulnerabilities, with interpretation success rate exceeding 90%.
Zhibin Jian, Siquan Huang, Hongyi Xie, Ying Gao 0004, Leyu Shi
Cybersecur.2
2026 Privacy-Preserving Rényi Layer-Wise Budget Allocation Against Gradient Leakage for Federated Learning
abstract
Federated learning (FL) is vulnerable to gradient-based privacy attacks, where malicious attackers reconstruct training data from exchanged gradients. While existing differential privacy (DP) defenses mitigate this, they often cause excessive additive noise due to the inequality scaling in the theoretical analyses, which degrades the model's utility or fail under adaptive attacks. To address this issue, we proposeFedMSBA, a layer-wise privacy-preservation method that adaptively allocates privacy budgets via Rényi DP (RDP) and modified sensitivity. FedMSBA dynamically scales noise to model intricacies and adaptively choose the better applied DP mechanisms, which provides a tighter mathematical bound and finally prevents non-convergence while resisting reconstruction attacks. Experiments demonstrate superior privacy-utility trade-offs compared to state-of-the-art defenses. FedMSBA achieves an approximately 2% improvement in accuracy and a 5% enhancement in privacy preservation. Furthermore, FedMSBA's performance remains nearly unaffected by variations in the privacy budget$\epsilon$and failure rate$\delta$.
Leyu Shi, Ying Gao 0004, Chong Chen 0011, Siquan Huang, Jiafeng Zhao, Xiping Hu
IEEE Trans. Mob. Comput.4
2025 Unsupervised Histopathological Image Semantic Segmentation with Overlapping Patches Consistency Constraint
Wentian Cai, Weizhao Weng, Yandan Chen, Siquan Huang, Victor C. M. Leung, Ying Gao 0004
ICCV5
2025 Distributed clustering meets federated learning: a clustering-based approach to data poisoning mitigation
abstract
Abstract Data poisoning attacks present a significant challenge to the integrity and reliability of federated learning (FL) systems, where model training occurs collaboratively across decentralized devices. These attacks involve the deliberate injection of malicious data to corrupt the model’s training process, ultimately undermining its performance. Given the decentralized nature of FL and the lack of direct access to local data, detecting and mitigating these attacks becomes particularly difficult, especially in unsupervised scenarios where labeled data is unavailable. In this paper, we introduce a novel Federated Data Sanitization Defense to address these security threats in federated learning environments. This defense mechanism leverages federated clustering to group model updates based on semantic consistency, identifying and isolating outlier updates that are likely to be poisoned. A targeted data sanitization strategy is then applied to filter out malicious data, ensuring that only trustworthy information is used to update the global model. This decentralized process occurs on each participating device, enabling real-time detection and mitigation of data poisoning attacks. Through extensive experiments, we validate the effectiveness of Federated Data Sanitization Defense, demonstrating its ability to enhance the security and robustness of federated learning systems against data poisoning, while preserving privacy and model integrity.
Chong Chen 0011, Siquan Huang, Leyu Shi, Ying Gao 0004
Cybersecur.2
2025 FedMAR: A Privacy-Preserving and Robust Server-Side Multistage Federated Learning
abstract
In recent years, federated learning (FL) has continued to evolve with the advent of big data and the large language model (LLM), but it has also exposed numerous security and privacy issues. As a form of distributed machine learning, FL systems are more susceptible to poisoning attacks because training data are dispersed across different participants; additionally, the training achievement of FL may be subject to low-cost theft by some free-riders. Existing works have addressed defenses against the aforementioned two types of threats, but they often focus on defending against only one type and fail to effectively integrate defenses against multiple types of threats. However, in real-world Internet of Things (IoT) systems, the types of threats are not limited to just one category. In this work, we try to maintain the performance of the global model under poisoning attacks, preserve the privacy of the server under free-riders, and explore the balance between these two aspects. Therefore, this work proposes Federated Multi-Stage Asynchronous Roll-back (FedMAR), ensuring the quality of local updates; in addition, this work also provides privacy preservation in the global update process based on Rinyi Differential Privacy (RDP), and offers a certain basis for detecting free-riders. To validate the generalization of the proposed method, we conducted relevant experiments on both image and text datasets, and further investigated the robustness of the proposed method against poisoning attacks, model inversion attacks, data heterogeneity, and other aspects. The testing accuracy of the global model can even be improved by 7.2%.
Leyu Shi, Ying Gao 0004, Chong Chen 0011, Siquan Huang, Jiafeng Zhao, Xiping Hu, Victor C. M. Leung
IEEE Internet Things J.4
2025 FedCleanse: Cleanse the backdoor attacks in federated learning system
Siquan Huang, Yijiang Li, Chong Chen 0011, Leyu Shi, Wentian Cai, Ying Gao 0004
Knowl. Based Syst.1
2025 FedID: Enhancing Federated Learning Security Through Dynamic Identification
abstract
Federated learning (FL), recognized for its decentralized and privacy-preserving nature, faces vulnerabilities to backdoor attacks that aim to manipulate the model's behavior on attacker-chosen inputs. Most existing defenses based on statistical differences take effect only against specific attacks. This limitation becomes significantly pronounced when malicious gradients closely resemble benign ones or the data exhibits non-IID characteristics, making the defenses ineffective against stealthy attacks. This paper revisits distance-based defense methods and uncovers two critical insights: First, Euclidean distance becomes meaningless in high dimensions. Second, a single metric cannot identify malicious gradients with diverse characteristics. As a remedy, we propose FedID, a simple yet effective strategy employing multiple metrics with dynamic weighting for adaptive backdoor detection. Besides, we present a modified z-score approach to select the gradients for aggregation. Notably, FedID does not rely on predefined assumptions about attack settings or data distributions and minimally impacts benign performance. We conduct extensive experiments on various datasets and attack scenarios to assess its effectiveness. FedID consistently outperforms previous defenses, particularly excelling in challenging Edge-case PGD scenarios. Our experiments highlight its robustness against adaptive attacks tailored to break the proposed defense and adaptability to a wide range of non-IID data distributions without compromising benign performance.
Siquan Huang, Yijiang Li, Chong Chen 0011, Ying Gao 0004, Xiping Hu
IEEE Trans. Pattern Anal. Mach. Intell.1
2025 Scope: On Detecting Constrained Backdoor Attacks in Federated Learning
abstract
Federated learning (FL) allows multiple clients to train an efficient deep-learning model collaboratively but is susceptible to backdoor attacks. Traditional detection-based defenses depend on specific metrics to distinguish client gradients. Defense-aware attackers exploit this by constraining attack gradients on these metrics to evade detection, leading to metric-constrained attacks. This paper concretely instantiates such threats and introduces cosine-constrained attacks, which successfully compromise advanced defenses based on cosine distance. To address the aforementioned challenge, we propose Scope, a novel defense that detects cosine-constrained attacks using cosine distance by exposing the constrained backdoor dimensions of attack gradients. Scope employs dimension-wise normalization and differential scaling to amplify the distinction between backdoor dimensions and benign or unused ones, countering sophisticated attackers’ attempts to obscure them. Moreover, we develop a novel clustering approach, namely Dominant Gradient Clustering (DGC), to isolate and eliminate backdoor gradients. Extensive experiments across various datasets, models, FL settings, and adversary scenarios demonstrate that Scope consistently outperforms existing defenses by a significant margin, especially against the cosine-constrained attack. Additionally, we present a Scope-tailored attack designed to evade Scope, but it remains ineffective even when maximizing stealthiness, further underscoring the robustness of Scope. We release our source code at:https://github.com/siquanhuang/Scope.
Siquan Huang, Yijiang Li, Xingfu Yan, Ying Gao 0004, Chong Chen 0011, Leyu Shi, Wing W. Y. Ng
IEEE Trans. Inf. Forensics Secur.1
2024 GN: Guided Noise Eliminating Backdoors in Federated Learning
abstract
Federated learning (FL) trains a model collaboratively but is susceptible to backdoor attacks for its privacy-preserving nature. Existing defenses against backdoor attacks in FL always make specific assumptions on data distributions among clients and are ineffective against sophisticated attacks. Although adding noise mitigates backdoors injected in the model, it simultaneously negatively impacts the main performance. To address the aforementioned issues, we propose a novel defense mechanism, Guided Noise (GN), that eliminates backdoors without compromising the model's main performance. GN achieves this by utilizing conductance to evaluate the importance of neurons and subsequently adding guided noise to suspected backdoor neurons selected by voting, which only disturbs the backdoor task. Extensive experimental evaluations of GN show its significant superiority over traditional noising-based defenses, making it a valuable replacement for existing noising to enhance the robustness of existing defenses against backdoor attacks in FL.
Siquan Huang, Ying Gao 0004, Chong Chen 0011, Leyu Shi
SMC1
2023 Multi-metrics adaptively identifies backdoors in Federated learning
abstract
The decentralized and privacy-preserving nature of federated learning (FL) makes it vulnerable to backdoor attacks aiming to manipulate the behavior of the resulting model on specific adversary-chosen inputs. However, most existing defenses based on statistical differences take effect only against specific attacks, especially when the malicious gradients are similar to benign ones or the data are highly non-independent and identically distributed (non-IID). In this paper, we revisit the distance-based defense methods and discover that i) Euclidean distance becomes meaningless in high dimensions and ii) malicious gradients with diverse characteristics cannot be identified by a single metric. To this end, we present a simple yet effective defense strategy with multi-metrics and dynamic weighting to identify backdoors adaptively. Furthermore, our novel defense has no reliance on predefined assumptions over attack settings or data distributions and little impact on benign performance. To evaluate the effectiveness of our approach, we conduct comprehensive experiments on different datasets under various attack settings, where our method achieves the best defensive performance. For instance, we achieve the lowest backdoor accuracy of 3.06% under the most difficult Edge-case PGD, showing significant superiority over previous defenses. The experiments also demonstrate that our method can be well-adapted to a wide range of non-IID degrees without sacrificing the benign performance.
Siquan Huang, Yijiang Li, Chong Chen 0011, Leyu Shi, Ying Gao 0004
ICCV1