VLDB 2026 Research / reviewers in the wild / expert
Binghui Wang
dblp:123/7149
· DBLP profile ↗
84ranked-venue papers
19as first author
67since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 41 · 11 first-author · 35 since 2021Security and privacy · 29 · 5 first-author · 23 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 4 first-author · 17 since 2021Databases, data management, data science and information retrieval · 12 · 5 first-author · 8 since 2021Systems, architecture and hardware · 4 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 since 2021Computer networks · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Large-Scale Security Analysis of Multi-Token Smart Contracts: Uncovering Hidden Flaws in Batch Transfers
Ashok Kasthuri, Sajad Meisami, Lingxiao Jiang, Binghui Wang, Yue Duan |
DSN | 4 |
| 2026 | Ensemble Conformal Predictor (EnCP): A New Conformal Predictor with Robustness Guarantees Against Data Poisoning Attacks
Runyang Feng, Liren Shan, Binghui Wang |
SP | 5 |
| 2026 | IPG-FRN: Intrinsic prototype-guided feature reconstruction network for industrial anomaly detection
Lanxiao Li, Yishuo Liu, Binghui Wang |
Expert Syst. Appl. | 4 |
| 2026 | Prompt-guided fine-grained semantic alignment for weakly supervised video anomaly detectionabstractWeakly-supervised Video Anomaly Detection (wVAD) detects frame-level anomalies using only coarse video labels. Binary video-level supervision makes it hard for models to pinpoint exact anomaly boundaries. Most methods also rely solely on visual features, making it difficult to distinguish visually similar but semantically different events (e.g., hugging vs. fighting), leading to false alarms. To address these issues, we propose the P rompt- G uided Fine-Grained S emantic A lignment for Weakly-supervised V ideo A nomaly D etection ( PGSA-VAD ), which uses prompt-based priors to enhance anomaly distinction under weak supervision. Specifically, we propose the D ynamic M agnitude- D riven T emporal A ggregation M odule ( DMD-TAM ), which adaptively extracts local anomaly features and aggregates multi-scale fine-grained feature aggregation, achieving comprehensive temporal modeling while maintaining computational efficiency. Additionally, our P rompt- G uided F eature E nhancement M odule ( PGFEM ) combines textural prior knowledge with visual features to improve class separability and clarify anomaly boundaries. Extensive experiments demonstrate the effectiveness of PGSA-VAD on three public datasets. Binghui Wang, Chunjuan Yan, Lanxiao Li |
Knowl. Based Syst. | 1 |
| 2026 | A prototype correction multi-scale feature reconstruction network for industrial anomaly detection
Lanxiao Li, Chunjuan Yan, Binghui Wang |
Pattern Recognit. | 4 |
| 2026 | Self-supervised learning video anomaly detection based on time interval prediction and noise classification
Yishuo Liu, Qingyang Yang, Lanxiao Li, Binghui Wang |
Pattern Recognit. | 5 |
| 2026 | GMSFormer: Geometric-aware multi-structured transformer for point cloud understanding
Chunjuan Yan, Binghui Wang, Yishuo Liu |
Pattern Recognit. | 3 |
| 2026 | A Tri-Factor Adaptive Federated Learning Framework for Parkinson's Disease Diagnosis via Multi-Source Facial Expression AnalysisabstractEarly diagnosis of Parkinson's disease (PD) is crucial for timely treatment and disease management. Recent studies link PD to impaired facial muscle control, manifesting as "masked face" symptoms, offering a novel diagnostic approach through facial expression analysis. However, data privacy concerns and legal restrictions have resulted in significant "data silos", hindering data sharing and limiting the accuracy and generalizability of existing diagnostic models due to small, localized datasets. To address these challenges, we propose an innovative Tri-Factor Adaptive Federated Learning (TriAFL) framework, designed to collaboratively analyze facial expression data across multiple medical institutions while ensuring robust data privacy protection. TriAFL introduces a comprehensive evaluation mechanism that assesses client contributions across three dimensions: gradient, data, and learning efficiency, effectively addressing Non-IID issues arising from data size variations and heterogeneity. To validate the real-world applicability of our method, we collaborate with a hospital to build the largest known facial expression dataset of PD patients. Furthermore, we explore the integration of local data augmentation strategy to further enhance diagnostic accuracy. Comprehensive experimental results demonstrate TriAFL's superior performance over conventional FL methods in classification task, as well as confirms TriAFL's efficacy in PD diagnosis, delivering a rapid, non-invasive screening tool while driving advancements in AI-powered healthcare. Houwei Xu, Yintao Zhou, Shengbo Chen, Binghui Wang, Wei Huang 0013 |
IEEE J. Biomed. Health Informatics | 6 |
| 2025 | Practicable Black-Box Evasion Attacks on Link Prediction in Dynamic Graphs - a Graph Sequential Embedding MethodabstractLink prediction in dynamic graphs (LPDG) has been widely applied to real-world applications such as website recommendation, traffic flow prediction, organizational studies, etc. These models are usually kept local and secure, with only the interactive interface restrictively available to the public. Thus, the problem of the black-box evasion attack on the LPDG model, where model interactions and data perturbations are restricted, seems to be essential and meaningful in practice. In this paper, we propose the first practicable black-box evasion attack method that achieves effective attacks against the target LPDG model, within a limited amount of interactions and perturbations. To perform effective attacks under limited perturbations, we develop a graph sequential embedding model to find the desired state embedding of the dynamic graph sequences, under a deep reinforcement learning framework. To overcome the scarcity of interactions, we design a multi-environment training pipeline and train our agent for multiple instances, by sharing an aggregate interaction buffer. Finally, we evaluate our attack against three advanced LPDG models on three real-world graph datasets of different scales and compare its performance with related methods under the interaction and perturbation constraints. Experimental results show that our attack is both effective and practicable. Jiate Li, Binghui Wang |
AAAI | 3 |
| 2025 | Breaking Data Silos in Parkinson's Disease Diagnosis: An Adaptive Federated Learning Approach for Privacy-Preserving Facial Expression AnalysisabstractThe early diagnosis of Parkinson’s disease (PD) is crucial for potential patients to receive timely treatment and prevent disease progression. Recent studies have shown that PD is closely linked to impairments in facial muscle control, resulting in characteristic “masked face” symptoms. This discovery offers a novel perspective for PD diagnosis by leveraging facial expression recognition and analysis techniques to capture and quantify these features, thereby distinguishing between PD patients and non-PD individuals based on their facial expressions. However, concerns about data privacy and legal restrictions have led to significant “data silos”, posing challenges to data sharing and limiting the accuracy and generalization of existing diagnostic models due to small, localized datasets. To address this issue, we propose an innovative adaptive federated learning approach that aims to jointly analyze facial expression data from multiple medical institutions while preserving data privacy. Our proposed approach comprehensively evaluates each client's contributions in terms of gradient, data, and learning efficiency, overcoming the non-IID issues caused by varying data sizes or heterogeneity across clients. To demonstrate the real-world impact of our approach, we collected a new facial expression dataset of PD patients in collaboration with a hospital. Extensive experiments validate the effectiveness of our proposed method for PD diagnosis and facial expression recognition, offering a promising avenue for rapid, non-invasive initial screening and advancing healthcare intelligence. Houwei Xu, Yintao Zhou, Wei Huang 0013, Binghui Wang |
AAAI | 6 |
| 2025 | Learning Robust and Privacy-Preserving Representations via Information TheoryabstractMachine learning models are vulnerable to both security attacks (e.g., adversarial examples) and privacy attacks (e.g., private attribute inference). We take the first step to mitigate both the security and privacy attacks, and maintain task utility as well. Particularly, we propose an information-theoretic framework to achieve the goals through the lens of representation learning, i.e., learning representations that are robust to both adversarial examples and attribute inference adversaries. We also derive novel theoretical results under our framework, e.g., the inherent trade-off between adversarial robustness/utility and attribute privacy, and guaranteed attribute privacy leakage against attribute inference adversaries. Binghui Zhang, Sayedeh Leila Noorbakhsh, Yuan Hong 0001, Binghui Wang |
AAAI | 5 |
| 2025 | Harmonizing Differential Privacy Mechanisms for Federated Learning: Boosting Accuracy and ConvergenceabstractDifferentially private federated learning (DP-FL) offers a compelling approach to collaborative model training by ensuring robust privacy for clients. Despite its potential, current methods face challenges in effectively balancing privacy, utility, and performance across diverse federated learning scenarios. Addressing these challenges, we introduce UDP-FL, to our knowledge the first DP-FL framework that universally harmonizes any randomization mechanism, including those considered optimal, by employing the Gaussian Moments Accountant (viz. DP-SGD). Central to UDP-FL is the 'Harmonizer,' a dynamic module engineered to intelligently select and apply the most suitable DP mechanism tailored to each client's specific privacy requirements, data sensitivities, and computational capacities. This selection process is driven by the principle of Rényi Differential Privacy, which serves as a crucial mediator for aligning privacy budgets effectively. Our comprehensive evaluation of UDP-FL, benchmarked against established baseline methods, demonstrates superior performance in upholding privacy guarantees and enhancing model functionality. The framework's robustness has been rigorously tested against a broad spectrum of privacy attacks, making it one of the most thorough validations of a DP-FL framework to date. Shuya Feng, Meisam Mohammady, Hanbin Hong, Shenao Yan, Ashish Kundu, Binghui Wang, Yuan Hong 0001 |
CODASPY | 6 |
| 2025 | Deterministic Certification of Graph Neural Networks against Graph Poisoning Attacks with Arbitrary PerturbationsabstractGraph neural networks (GNNs) are becoming the de facto method to learn on the graph data and have achieved the state-of-the-art on node and graph classification tasks. However, recent works show GNNs are vulnerable to training-time poisoning attacks – marginally perturbing edges, nodes, or/and node features of training graph(s) can largely degrade GNNs’ testing performance. Most previous defenses against graph poisoning attacks are empirical and are soon broken by adaptive / stronger ones. A few provable defenses provide robustness guarantees, but have large gaps when applied in practice: 1) restrict the attacker on only one type of perturbation; 2) design for a particular GNN architecture or task; and 3) robustness guarantees are not 100% accurate.In this work, we bridge all these gaps by developing PGNNCert, the first certified defense of GNNs against poisoning attacks under arbitrary (edge, node, and node feature) perturbations with deterministic robustness guarantees. Extensive evaluations on multiple node and graph classification datasets and GNNs demonstrate the effectiveness of PGNNCert to provably defend against arbitrary poisoning perturbations. PGNNCert is also shown to significantly outperform the state-of-the-art certified defenses against edge perturbation or node perturbation during GNN training. Jiate Li, Binghui Wang |
CVPR | 4 |
| 2025 | Provably Robust Explainable Graph Neural Networks against Graph Perturbation AttacksabstractExplaining Graph Neural Network (XGNN) has gained growing attention to facilitate the trust of using GNNs, which is the mainstream method to learn graph data. Despite their growing attention, Existing XGNNs focus on improving the explanation performance, and its robustness under attacks is largely unexplored. We noticed that an adversary can slightly perturb the graph structure such that the explanation result of XGNNs is largely changed. Such vulnerability of XGNNs could cause serious issues particularly in safety/security-critical applications. In this paper, we take the first step to study the robustness of XGNN against graph perturbation attacks, and propose XGNNCert, the first provably robust XGNN. Particularly, our XGNNCert can provably ensure the explanation result for a graph under the worst-case graph perturbation attack is close to that without the attack, while not affecting the GNN prediction, when the number of perturbed edges is bounded. Evaluation results on multiple graph datasets and GNN explainers show the effectiveness of XGNNCert. Jiate Li, Jinyuan Jia 0001, Binghui Wang |
ICLR | 5 |
| 2025 | Measure-Theoretic Anti-Causal Representation LearningabstractCausal representation learning in the anti-causal setting—labels cause features rather than the reverse—presents unique challenges requiring specialized approaches. We propose Anti-Causal Invariant Abstractions (ACIA), a novel measure-theoretic framework for anti-causal representation learning. ACIA employs a two-level design: low-level representations capture how labels generate observations, while high-level representations learn stable causal patterns across environment-specific variations. ACIA addresses key limitations of existing approaches by: (1) accommodating prefect and imperfect interventions through interventional kernels, (2) eliminating dependency on explicit causal structures, (3) handling high-dimensional data effectively, and (4) providing theoretical guarantees for out-of-distribution generalization. Experiments on synthetic and real-world medical datasets demonstrate that ACIA consistently outperforms state-of-the-art methods in both accuracy and invariance metrics. Furthermore, our theoretical results establish tight bounds on performance gaps between training and unseen environments, confirming the efficacy of our approach for robust anti-causal learning. {{Code is available at \url{https://github.com/ArmanBehnam/ACIA}}}. Arman Behnam, Binghui Wang |
NeurIPS | 2 |
| 2025 | AGNNCert: Defending Graph Neural Networks against Arbitrary Perturbations with Deterministic Certification
Jiate Li, Binghui Wang |
USENIX Security Symposium | 2 |
| 2025 | Rectifying Privacy and Efficacy Measurements in Machine Unlearning: A New Inference Attack Perspective
Nima Naderloui, Shenao Yan, Binghui Wang, Jie Fu 0003, Wendy Hui Wang, Yuan Hong 0001 |
USENIX Security Symposium | 3 |
| 2025 | PoisonedRAG: Knowledge Corruption Attacks to Retrieval-Augmented Generation of Large Language Models
Runpeng Geng, Binghui Wang, Jinyuan Jia 0001 |
USENIX Security Symposium | 3 |
| 2025 | MMI-FM: Multimodal Interaction and Fusion Mechanism for Video Anomaly Detection
Binghui Wang, Jiajiong Li, Yishuo Liu |
Appl. Intell. | 1 |
| 2025 | Recipient-Aware Photo Automatic Deletion Control Policy Recommendation Scheme in Online Social NetworksabstractContent sharing, whether in Online Social Networks (OSNs) or even in the Internet of Things (IoT), serves as a pivotal link in the flow of data. To better protect the privacy of shared content, current OSNs allow sharers to manually set policies for uploaded content. However, this method of policy setting is not suitable for scenarios where IoT is deeply integrated with OSNs, as IoT devices often share content frequently and automatically. To address this issue, we propose the design, implementation, and evaluation of SmartCircles, a personalized photo-sharing and automatic deletion scheme. SmartCircles can function as a plugin within existing OSNs, supporting operations on various smart devices. It encompasses the following steps: a) Before sharing a photo, calculate the intimacy level depicted in the photo and the sharer's willingness to share. b) Before the recipient views the photo, calculate the intimacy between the sharer and the recipient, and evaluate feedback from the recipient. c) Based on the results computed above and a trade-off between profit and loss, recommend a recipient-aware automatic deletion control policy for the photo. We implement a prototype of SmartCircles, and the evaluation results demonstrate its effectiveness with an accuracy rate of policy recommendations reaching approximately 92%. Haiyang Luo, Zhe Sun 0005, Yunqing Sun, Ang Li 0005, Binghui Wang, Jin Cao 0001, Ben Niu 0001 |
IEEE Trans. Dependable Secur. Comput. | 5 |
| 2025 | Identifying Backdoored Graphs in Graph Neural Network Training: An Explanation-Based Approach With Novel MetricsabstractGraph Neural Networks (GNNs) have gained popularity in numerous domains, yet they are vulnerable to backdoor attacks that can compromise their performance and ethical application. The detection of these attacks is crucial for maintaining the reliability and security of GNN classification tasks, but existing methods are often inflexible, relying on single metrics that fail to capture the full range of backdoor behaviors. Recognizing the challenge in detecting such intrusions, we devised a novel detection method that creatively leverages graph-level explanations. By extracting and transforming secondary outputs from GNN explanation mechanisms, we developed seven innovative metrics for effective detection of backdoor attacks on GNNs. Additionally, we develop an adaptive attack to rigorously evaluate our approach. We test our method on multiple benchmark datasets and examine its efficacy against various attack models. Our results show that our method can achieve high detection performance, marking a significant advancement in safeguarding GNNs against backdoor attacks. Jane Downer, Ren Wang 0008, Binghui Wang |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2025 | A Secret Sharing-Inspired Robust Distributed Backdoor Attack to Federated LearningabstractFederated Learning (FL) is vulnerable to backdoor attacks—especially distributed backdoor attacks (DBA) that are more persistent and stealthy than centralized backdoor attacks. However, we observe that the attack effectiveness of DBA can be largely reduced when encountering rebels, i.e., the agents promising to perform the attack but do not do so. To robustify DBAs, we present SSRDBA , a secret sharing-inspired robust DBA to FL. To be specific, given a same global trigger as DBA, SSRDBA carefully divides it into different shares based on secret sharing and exploits these shares to poison local data on malicious devices, respectively. SSRDBA enjoys several merits, e.g., only partial malicious agents guarantee the reconstruction of the global trigger. Extensive experimental results show that SSRDBA is more robust to rebels than DBA and can evade the state-of-the-art FL defenses mainly for centralized backdoor attacks. To mitigate SSRDBA , we further design a novel defense mechanism, termed NFDR, which shows great potential against SSRDBA on certain independent identically distributed datasets. Yuxin Yang 0003, Qiang Li 0008, Yuede Ji, Binghui Wang |
ACM Trans. Priv. Secur. | 4 |
| 2025 | Align and Blend: A Unified Multi-Modal LiDAR Segmentation NetworkabstractThe multimodal fusion of point cloud data is a challenging task in 3D computer vision, particularly in fine object segmentation. In this study, we propose an innovative unified multimodal LiDAR segmentation network, named Align and Blend (A2Blend for short), which skillfully integrates three representative forms of point clouds: point view, voxel view, and range view. The core innovation of A2Blend lies in addressing two primary tasks: Align and Blend. For the “Align” task, we have carefully designed a learnable cross-modal association module, whose core is a unique “Cross-Modal Triplet Alignment Loss” mechanism. To the best of our knowledge, this is the first application of such a mechanism in the field of LiDAR segmentation. This mechanism applies principles from contrastive learning, promoting the tight clustering of similar semantic groups both within and across modalities, while increasing the separation between dissimilar groups by expanding the distance between similar and dissimilar sample clusters in the feature space. This significantly enhances the discriminative power and representational capacity of the feature embeddings. For the “Blend” task, we propose a fusion strategy that integrates intragroup self-attention within modalities and intergroup self-attention across modalities. This approach combines key concepts from standard self-attention and cross-self-attention mechanisms to achieve more comprehensive multimodal fusion. The segmentation performance on two large-scale outdoor datasets and one indoor dataset surpasses that of most state-of-the-art algorithms. Jiajiong Li, Binghui Wang |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2025 | Heterogeneous Prototype Learning From Contaminated Faces Across Domains via Disentangling Latent FactorsabstractThis article studies an emerging practical problem called heterogeneous prototype learning (HPL). Unlike the conventional heterogeneous face synthesis (HFS) problem that focuses on precisely translating a face image from a source domain to another target one without removing facial variations, HPL aims at learning the variation-free prototype of an image in the target domain while preserving the identity characteristics. HPL is a compounded problem involving two cross-coupled subproblems, that is, domain transfer and prototype learning (PL), thus making most of the existing HFS methods that simply transfer the domain style of images unsuitable for HPL. To tackle HPL, we advocate disentangling the prototype and domain factors in their respective latent feature spaces and then replacing the source domain with the target one for generating a new heterogeneous prototype. In doing so, the two subproblems in HPL can be solved jointly in a unified manner. Based on this, we propose a disentangled HPL framework, dubbed DisHPL, which is composed of one encoder-decoder generator and two discriminators. The generator and discriminators play adversarial games such that the generator embeds contaminated images into a prototype feature space only capturing identity information and a domain-specific feature space, while generating realistic-looking heterogeneous prototypes. Experiments on various heterogeneous datasets with diverse variations validate the superiority of DisHPL. Binghui Wang, Mang Ye, Yiu-Ming Cheung, Yintao Zhou, Wei Huang 0013, Bihan Wen |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | Task-Agnostic Privacy-Preserving Representation Learning for Federated Learning against Attribute Inference AttacksabstractFederated learning (FL) has been widely studied recently due to its property to collaboratively train data from different devices without sharing the raw data. Nevertheless, recent studies show that an adversary can still be possible to infer private information about devices' data, e.g., sensitive attributes such as income, race, and sexual orientation. To mitigate the attribute inference attacks, various existing privacy-preserving FL methods can be adopted/adapted. However, all these existing methods have key limitations: they need to know the FL task in advance, or have intolerable computational overheads or utility losses, or do not have provable privacy guarantees. We address these issues and design a task-agnostic privacy-preserving presentation learning method for FL (TAPPFL) against attribute inference attacks. TAPPFL is formulated via information theory. Specifically, TAPPFL has two mutual information goals, where one goal learns task-agnostic data representations that contain the least information about the private attribute in each device's data, and the other goal ensures the learnt data representations include as much information as possible about the device data to maintain FL utility. We also derive privacy guarantees of TAPPFL against worst-case attribute inference attacks, as well as the inherent tradeoff between utility preservation and privacy protection. Extensive results on multiple datasets and applications validate the effectiveness of TAPPFL to protect data privacy, maintain the FL utility, and be efficient as well. Experimental results also show that TAPPFL outperforms the existing defenses. Caridad Arroyo Arevalo, Sayedeh Leila Noorbakhsh, Yuan Hong 0001, Binghui Wang |
AAAI | 5 |
| 2024 | Certifiable Black-Box Attacks with Randomized Adversarial Examples: Breaking Defenses with Provable ConfidenceabstractBlack-box adversarial attacks have demonstrated strong potential to compromise machine learning models by iteratively querying the target model or leveraging transferability from a local surrogate model.Recently, such attacks can be effectively mitigated by state-of-the-art (SOTA) defenses, e.g., detection via the pattern of sequential queries, or injecting noise into the model. To our best knowledge, we take the first step to study a new paradigm of black-box attacks with provable guarantees -- certifiable black-box attacks that can guarantee the attack success probability (ASP) of adversarial examples before querying over the target model. This new black-box attack unveils significant vulnerabilities of machine learning models, compared to traditional empirical black-box attacks, e.g., breaking strong SOTA defenses with provable confidence, constructing a space of (infinite) adversarial examples with high ASP, and the ASP of the generated adversarial examples is theoretically guaranteed without verification/queries over the target model. Specifically, we establish a novel theoretical foundation for ensuring the ASP of the black-box attack with randomized adversarial examples (AEs). Then, we propose several novel techniques to craft the randomized AEs while reducing the perturbation size for better imperceptibility. Finally, we have comprehensively evaluated the certifiable black-box attacks on the CIFAR10/100, ImageNet, and LibriSpeech datasets, while benchmarking with 16 SOTA black-box attacks, against various SOTA defenses in the domains of computer vision and speech recognition. Both theoretical and experimental results have validated the significance of the proposed attack. Hanbin Hong, Xinyu Zhang 0016, Binghui Wang, Zhongjie Ba, Yuan Hong 0001 |
CCS | 3 |
| 2024 | Distributed Backdoor Attacks on Federated Graph Learning and Certified DefensesabstractFederated graph learning (FedGL) is an emerging federated learning (FL) framework that extends FL to learn graph data from diverse sources without accessing the data. FL for non-graph data has shown to be vulnerable to backdoor attacks, which inject a shared backdoor trigger into the training data such that the trained backdoored FL model can predict the testing data containing the trigger as the attacker desires. However, FedGL against backdoor attacks is largely unexplored, and no effective defense exists. Yuxin Yang 0003, Qiang Li 0008, Jinyuan Jia 0001, Yuan Hong 0001, Binghui Wang |
CCS | 5 |
| 2024 | Breaking State-of-the-Art Poisoning Defenses to Federated Learning: An Optimization-Based Attack FrameworkabstractFederated Learning (FL) is a novel client-server distributed learning framework that can protect data privacy. However, recent works show that FL is vulnerable to poisoning attacks. Many defenses with robust aggregators (AGRs) are proposed to mitigate the issue, but they are all broken by advanced attacks. Very recently, some renewed robust AGRs are designed, typically with novel clipping or/and filtering strategies, and they show promising defense performance against the advanced poisoning attacks. In this paper, we show that these novel robust AGRs are also vulnerable to carefully designed poisoning attacks. Specifically, we observe that breaking these robust AGRs reduces to bypassing the clipping or/and filtering of malicious clients, and propose an optimization-based attack framework to leverage this observation. Under the framework, we then design the customized attack against each robust AGR. Extensive experiments on multiple datasets and threat models verify our proposed optimizationbased attack can break the SOTA AGRs. We hence call for novel defenses against poisoning attacks to FL. Code is available at: https: //github.com/Yuxin104/BreakSTOAPoisoningDefenses. Yuxin Yang 0003, Qiang Li 0008, Chenfei Nie, Yuan Hong 0001, Binghui Wang |
CIKM | 5 |
| 2024 | Leveraging Local Structure for Improving Model Explanations: An Information Propagation ApproachabstractNumerous explanation methods have been recently developed to interpret the decisions made by deep neural network (DNN) models. For image classifiers, these methods typically provide an attribution score to each pixel in the image to quantify its contribution to the prediction. However, most of these explanation methods appropriate attribution scores to pixels independently, even though both humans and DNNs make decisions by analyzing a set of closely related pixels simultaneously. Hence, the attribution score of a pixel should be evaluated jointly by considering itself and its structurally-similar pixels. We propose a method called IProp, which models each pixel's individual attribution score as a source of explanatory information and explains the image prediction through the dynamic propagation of information across all pixels. To formulate the information propagation, IProp adopts the Markov Reward Process, which guarantees convergence, and the final status indicates the desired pixels' attribution scores. Furthermore, IProp is compatible with any existing attribution-based explanation method. Extensive experiments on various explanation methods and DNN models verify that IProp significantly improves them on a variety of interpretability metrics. Ruo Yang, Binghui Wang, Mustafa Bilgic 0001 |
CIKM | 2 |
| 2024 | Graph Neural Network Causal Explanation via Neural Causal Models
Arman Behnam, Binghui Wang |
ECCV (61) | 2 |
| 2024 | Early Diagnosing Parkinson's Disease Via a Deep Learning Model Based on Augmented Facial Expression DataabstractIt is crucial to promptly diagnose potential Parkinson's disease (PD) patients in order to facilitate early treatment and prevent disease progression. In recent years, there has been growing interest in using facial expressions for in-vitro PD diagnosis due to the distinct "masked face" characteristics of PD patients and the cost-effectiveness of this approach. However, current facial expression-based PD diagnosis methods are hindered by limited training data on PD patients' facial expressions and weak prediction models. To address these issues, we propose a new PD diagnosis method that utilizes facial expression data augmentation and deep neural network prediction. Our approach involves two stages: 1) generating virtual facial expression images depicting six basic emotions (anger, disgust, fear, happiness, sadness, and surprise) through multi-domain adversarial learning to expand the original training data; 2) training a deep neural network prediction model using a combination of the augmented training data from PD patients and facial expression images of normal individuals from public datasets. Qualitative and quantitative experiments confirm the efficacy of our multi-domain adversarial learning-based facial expression synthesis and demonstrate the promising performance of our proposed approach for PD diagnosis. Yintao Zhou, Wei Huang 0013, Binghui Wang |
ICASSP | 4 |
| 2024 | GNNCert: Deterministic Certification of Graph Neural Networks against Adversarial PerturbationsabstractGraph classification, which aims to predict a label for a graph, has many real-world applications such as malware detection, fraud detection, and healthcare. However, many studies show an attacker could carefully perturb the structure and/or node features in a graph such that a graph classifier misclassifies the perturbed graph. Such vulnerability impedes the deployment of graph classification in security/safety-critical applications. Existing empirical defenses lack formal robustness guarantees and could be broken by adaptive or unknown attacks. Existing provable defenses have the following limitations: 1) they achieve sub-optimal robustness guarantees for graph structure perturbation, 2) they cannot provide robustness guarantees for arbitrarily node feature perturbations, 3) their robustness guarantees are probabilistic, meaning they could be incorrect with a non-zero probability, and 4) they incur large computation costs. We aim to address those limitations in this work. We propose GNNCert, a certified defense against both graph structure and node feature perturbations for graph classification. Our GNNCert provably predicts the same label for a graph when the number of perturbed edges and the number of nodes with perturbed features are bounded. Our results on 8 benchmark datasets show that GNNCert outperforms three state-of-the-art methods. Zaishuo Xia, Binghui Wang, Jinyuan Jia 0001 |
ICLR | 3 |
| 2024 | Reconstructing Prototype From Contaminated Face With Variations Across Heterogeneous DomainsabstractThis paper focuses on a new heterogeneous prototype learning (HPL) problem, which aims at reconstructing the variation-free and identity-preserved prototype in the target domain from a contaminated input image in the source domain. Most existing heterogeneous face synthesis (HFS) methods are unsuitable for HPL, as these methods focus on performing accurate image-to-image translation with facial details unaltered, but cannot effectively remove the input facial variations. In this paper, we propose an identity-aware cycle-consistent network, dubbed IAC2N, for image-to-prototype transformation across domains. To address HPL, IAC2N designs three effective losses, i.e., prototype adversarial loss, label information guided identity loss, and prototype learning cycle loss, in its objective. The first loss is used for transferring the domain style as well as removing the universal facial variations. The latter two losses are used for maintaining the identity consistency during HPL from an explicit and an implicit perspectives, respectively. Furthermore, IAC2N is a joint learning framework that is able to learn the identity feature for the contaminated image via its encoder-decoder structural generator in order to perform heterogeneous face recognition (HFR). Extensive experiments on various heterogeneous face datasets demonstrate the effectiveness of IAC2N in both tasks of HPL and HFR. Binghui Wang, Nanrun Zhou, Yintao Zhou, Wei Huang 0013 |
ICME | 2 |
| 2024 | Graph Neural Network Explanations are FragileabstractExplainable Graph Neural Network (GNN) has emerged recently to foster the trust of using GNNs. Existing GNN explainers are developed from various perspectives to enhance the explanation performance. We take the first step to study GNN explainers under adversarial attack—We found that an adversary slightly perturbing graph structure can ensure GNN model makes correct predictions, but the GNN explainer yields a drastically different explanation on the perturbed graph. Specifically, we first formulate the attack problem under a practical threat model (i.e., the adversary has limited knowledge about the GNN explainer and a restricted perturbation budget). We then design two methods (i.e., one is loss-based and the other is deduction-based) to realize the attack. We evaluate our attacks on various GNN explainers and the results show these explainers are fragile. Jiate Li, Jinyuan Jia 0001, Binghui Wang |
ICML | 5 |
| 2024 | FedGMark: Certifiably Robust Watermarking for Federated Graph LearningabstractFederated graph learning (FedGL) is an emerging learning paradigm to collaboratively train graph data from various clients. However, during the development and deployment of FedGL models, they are susceptible to illegal copying and model theft. Backdoor-based watermarking is a well-known method for mitigating these attacks, as it offers ownership verification to the model owner. We take the first step to protect the ownership of FedGL models via backdoor-based watermarking. Existing techniques have challenges in achieving the goal: 1) they either cannot be directly applied or yield unsatisfactory performance; 2) they are vulnerable to watermark removal attacks; and 3) they lack of formal guarantees. To address all the challenges, we propose FedGMark, the first certified robust backdoor-based watermarking for FedGL. FedGMark leverages the unique graph structure and client information in FedGL to learn customized and diverse watermarks. It also designs a novel GL architecture that facilitates defending against both the empirical and theoretically worst-case watermark removal attacks. Extensive experiments validate the promising empirical and provable watermarking performance of FedGMark. Source code is available at: https://github.com/Yuxin104/FedGMark. Yuxin Yang 0003, Qiang Li 0008, Yuan Hong 0001, Binghui Wang |
NeurIPS | 4 |
| 2024 | DeepTheft: Stealing DNN Model Architectures through Power Side ChannelabstractDeep Neural Network (DNN) models are often deployed in resource-sharing clouds as Machine Learning as a Service (MLaaS) to provide inference services. To steal model architectures that are of valuable intellectual properties, a class of attacks has been proposed via different side-channel leakage, posing a serious security challenge to MLaaS.Also targeting MLaaS, we propose a new end-to-end attack, DeepTheft, to accurately recover complex DNN model architectures on general processors via the RAPL (Running Average Power Limit)-based power side channel. While unprivileged access to the RAPL has been disabled in bare-metal OSes, we observe that the RAPL is still legitimately accessible in a platform as a service, e.g., the latest docker environment of version 20.10.18 used in this work. However, an attacker can acquire only a low sampling rate (1 KHz) of the time-series energy traces from the RAPL interface, rendering existing techniques ineffective in stealing large and deep DNN models. To this end, we design a novel and generic learning-based framework consisting of a set of meta-models, based on which DeepTheft is demonstrated to have high accuracy in recovering a large number (thousands) of models architectures from different model families including the deepest ResNet152. Particularly, DeepTheft has achieved a Levenshtein Distance Accuracy of 99.75% in recovering network structures, and a weighted average F1 score of 99.60% in recovering diverse layer-wise hyperparameters. Besides, our proposed learning framework is general to other time-series side-channel signals. To validate its generalization, another existing side channel is exploited, i.e., CPU frequency. Different from RAPL, CPU frequency is accessible to unprivileged users in bare-metal OSes. By using our generic learning framework trained against CPU frequency traces, DeepTheft has shown similarly high attack performance in stealing model architectures. Yansong Gao 0001, Huming Qiu, Zhi Zhang 0001, Binghui Wang, Alsharif Abuadbba, Minhui Xue 0001, Anmin Fu, Surya Nepal |
SP | 4 |
| 2024 | Text-CRS: A Generalized Certified Robustness Framework against Textual Adversarial AttacksabstractThe language models, especially the basic text classification models, have been shown to be susceptible to textual adversarial attacks such as synonym substitution and word insertion attacks. To defend against such attacks, a growing body of research has been devoted to improving the model’s robustness. However, providing provable robustness guarantees instead of empirical robustness is still widely unexplored. In this paper, we propose Text-CRS, a generalized certified robustness framework for natural language processing (NLP) based on randomized smoothing. To our best knowledge, existing certified schemes for NLP can only certify the robustness against ℓ0perturbations in synonym substitution attacks. Representing each word-level adversarial operation (i.e., synonym substitution, word reordering, insertion, and deletion) as a combination of permutation and embedding transformation, we propose novel smoothing theorems to derive robustness bounds in both permutation and embedding space against such adversarial operations. To further improve certified accuracy and radius, we consider the numerical relationships between discrete words and select proper noise distributions for the randomized smoothing. Finally, we conduct substantial experiments on multiple language models and datasets. Text-CRS can address all four different word-level adversarial operations and achieve a significant accuracy improvement. We also provide the first benchmark on certified accuracy and radius of four word-level operations, besides outperforming the state-of-the-art certification against synonym substitution attacks.1 Xinyu Zhang 0016, Hanbin Hong, Yuan Hong 0001, Binghui Wang, Zhongjie Ba, Kui Ren 0001 |
SP | 5 |
| 2024 | Inf2Guard: An Information-Theoretic Framework for Learning Privacy-Preserving Representations against Inference Attacks
Sayedeh Leila Noorbakhsh, Binghui Zhang, Yuan Hong 0001, Binghui Wang |
USENIX Security Symposium | 4 |
| 2024 | Efficient, Direct, and Restricted Black-Box Graph Evasion Attacks to Any-Layer Graph Neural Networks via Influence FunctionabstractGraph neural network (GNN), the mainstream method to learn on graph data, is vulnerable to graph evasion attacks, where an attacker slightly perturbing the graph structure can fool trained GNN models. Existing work has at least one of the following drawbacks: 1) limited to directly attack two-layer GNNs; 2) inefficient; and 3) impractical, as they need to know full or part of GNN model parameters. Binghui Wang, Minhua Lin, Tianxiang Zhou, Pan Zhou 0001, Ang Li 0005, Hai Li 0001, Yiran Chen 0001 |
WSDM | 1 |
| 2023 | Turning Strengths into Weaknesses: A Certified Robustness Inspired Attack Framework against Graph Neural NetworksabstractGraph neural networks (GNNs) have achieved state-of-the-art performance in many graph learning tasks. However, recent studies show that GNNs are vulnerable to both test-time evasion and training-time poisoning attacks that perturb the graph structure. While existing attack methods have shown promising attack performance, we would like to design an attack framework to further enhance the performance. In particular, our attack framework is inspired by certified robustness, which was originally used by defenders to defend against adversarial attacks. We are the first, from the attacker perspective, to leverage its properties to better attack GNNs. Specifically, we first derive nodes' certified perturbation sizes against graph evasion and poisoning attacks based on randomized smoothing, respectively. A larger certified perturbation size of a node indicates this node is theoretically more robust to graph perturbations. Such a property motivates us to focus more on nodes with smaller certified perturbation sizes, as they are easier to be attacked after graph perturbations. Accordingly, we design a certified robustness inspired attack loss, when incorporated into (any) existing attacks, produces our certified robustness inspired attack counterpart. We apply our frame-work to the existing attacks and results show it can significantly enhance the existing base attacks' performance. Binghui Wang |
CVPR | 1 |
| 2023 | IDGI: A Framework to Eliminate Explanation Noise from Integrated GradientsabstractIntegrated Gradients (IG) as well as its variants are well-known techniques for interpreting the decisions of deep neural networks. While IG-based approaches attain state-of-the-art performance, they often integrate noise into their explanation saliency maps, which reduce their interpretability. To minimize the noise, we examine the source of the noise analytically and propose a new approach to reduce the explanation noise based on our analytical findings. We propose the Important Direction Gradient Integration (IDGI) framework, which can be easily incorporated into any IG-based method that uses the Reimann Integration for integrated gradient computation. Extensive experiments with three IG-based methods show that IDGI improves them drastically on numerous interpretability metrics. The source code for IDGI is available at https://github.com/yangruo1226/IDGI. Ruo Yang, Binghui Wang, Mustafa Bilgic 0001 |
CVPR | 2 |
| 2023 | A Certified Radius-Guided Attack Framework to Image Segmentation ModelsabstractImage segmentation is an important problem in many safety-critical applications such as medical imaging and autonomous driving. Recent studies show that modern image segmentation models are vulnerable to adversarial perturbations, while existing attack methods mainly follow the idea of attacking image classification models. We argue that image segmentation and classification have inherent differences, and design an attack framework specially for image segmentation models. Our goal is to thoroughly explore the vulnerabilities of modern segmentation models, i.e., aiming to misclassify as many pixels as possible under a perturbation budget in both white-box and black-box settings.Our attack framework is inspired by certified radius, which was originally used by defenders to defend against adversarial perturbations to classification models. We are the first, from the attacker perspective, to leverage the properties of certified radius and propose a certified radius guided attack framework against image segmentation models. Specifically, we first adapt randomized smoothing, the state-of-the-art certification method for classification models, to derive the pixel’s certified radius. A larger certified radius of a pixel means the pixel is theoretically more robust to adversarial perturbations. This observation inspires us to focus more on disrupting pixels with relatively smaller certified radii. Accordingly, we design a pixel-wise certified radius guided loss, when plugged into any existing white-box attack, yields our certified radius-guided white-box attack.Next, we propose the first black-box attack to image segmentation models via bandit. A key challenge is no gradient information is available. To address it, we design a novel gradient estimator, based on bandit feedback, which is query-efficient and provably unbiased and stable. We use this gradient estimator to design a projected bandit gradient descent (PBGD) attack. We further use pixels’ certified radii and design a certified radius-guided PBGD (CR-PBGD) attack. We prove our PBGD and CR-PBGD attacks can achieve asymptotically optimal attack performance with an optimal rate. We evaluate our certified-radius guided white-box and black-box attacks on multiple modern image segmentation models and datasets. Our results validate the effectiveness of our certified radius-guided attack framework. Wenjie Qu 0001, Youqi Li, Binghui Wang |
EuroS&P | 3 |
| 2023 | Interpreting Disparate Privacy-Utility Tradeoff in Adversarial Learning via Attribute CorrelationabstractAdversarial learning is commonly used to extract latent data representations which are expressive to predict the target attribute but indistinguishable in the privacy attribute. However, whether they can achieve an expected privacy-utility tradeoff is of great uncertainty. In this paper, we posit it is the complex interaction between different attributes in the training set that causes disparate tradeoff results. We first formulate the measurement of utility, privacy and their tradeoff in adversarial learning. Then we propose the metrics of Statistical Reliability (SR) and Feature Reliability (FR) to quantify the relationship between attributes. Specifically, SR reflects the co-occurrence sampling bias of the joint distribution between two attributes. Beyond the explicit dependence, FR exploits the intrinsic interaction one attribute exerts on the other via exploring the representation disentanglement. We validate the metrics on CelebA and LFW dataset with a suite of target-privacy attribute pairs. Experimental results demonstrate the strong correlations between the metrics and utility, privacy and their tradeoff. We further conclude how to use SR and FR as a guide to the setting of the privacy-utility tradeoff parameter. Yahong Chen, Ang Li 0005, Binghui Wang, Yiran Chen 0001, Fenghua Li 0001, Jin Cao 0001, Ben Niu 0001 |
WACV | 4 |
| 2023 | DisP+V: A Unified Framework for Disentangling Prototype and Variation From Single Sample per PersonabstractSingle sample per person face recognition (SSPP FR) is one of the most challenging problems in FR due to the extreme lack of enrolment data. To date, the most popular SSPP FR methods are the generic learning methods, which recognize query face images based on the so-called prototype plus variation (i.e., P+V) model. However, the classic P+V model suffers from two major limitations: 1) it linearly combines the prototype and variation images in the observational pixel-spatial space and cannot generalize to multiple nonlinear variations, e.g., poses, which are common in face images and 2) it would be severely impaired once the enrolment face images are contaminated by nuisance variations. To address the two limitations, it is desirable to disentangle the prototype and variation in a latent feature space and to manipulate the images in a semantic manner. To this end, we propose a novel disentangled prototype plus variation model, dubbed DisP+V, which consists of an encoder-decoder generator and two discriminators. The generator and discriminators play two adversarial games such that the generator nonlinearly encodes the images into a latent semantic space, where the more discriminative prototype feature and the less discriminative variation feature are disentangled. Meanwhile, the prototype and variation features can guide the generator to generate an identity-preserved prototype and the corresponding variation, respectively. Experiments on various real-world face datasets demonstrate the superiority of our DisP+V model over the classic P+V model for SSPP FR. Furthermore, DisP+V demonstrates its unique characteristics in both prototype recovery and face editing/interpolation. Binghui Wang, Mang Ye, Yiu-Ming Cheung, Yiran Chen 0001, Bihan Wen |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2022 | NeuGuard: Lightweight Neuron-Guided Defense against Membership Inference AttacksabstractMembership inference attacks (MIAs) against machine learning models lead to serious privacy risks for the training dataset used in the model training. The state-of-the-art defenses against MIAs often suffer from poor privacy-utility balance and defense generality, as well as high training or inference overhead. To overcome these limitations, in this paper, we propose a novel, lightweight and effective Neuron-Guided Defense method named NeuGuard against MIAs. Unlike existing solutions which either regularize all model parameters in training or noise model output per input in real-time inference, NeuGuard aims to wisely guide the model output of training set and testing set to have close distributions through a fine-grained neuron regularization. That is, restricting the activation of output neurons and inner neurons in each layer simultaneously by using our developed class-wise variance minimization and layer-wise balanced output control. We evaluate NeuGuard and compare it with state-of-the-art defenses against two neural network based MIAs, five strongest metric based MIAs including the newly proposed label-only MIA on three benchmark datasets. Extensive experimental results show that NeuGuard outperforms the state-of-the-art defenses by offering much improved utility-privacy trade-off, generality, and overhead. Our code is publicly available at https://github.com/nux219/NeuGuard. Nuo Xu 0013, Binghui Wang, Wujie Wen, Parv Venkitasubramaniam |
ACSAC | 2 |
| 2022 | GraphTrack: A Graph-based Cross-Device Tracking FrameworkabstractCross-device tracking has drawn growing attention from both commercial companies and the general public because of its privacy implications and applications for user profiling, personalized services, etc. One particular, wide-used type of cross-device tracking is to leverage browsing histories of user devices, e.g., characterized by a list of IP addresses used by the devices and domains visited by the devices. However, existing browsing history based methods have three drawbacks. First, they cannot capture latent correlations among IPs and domains. Second, their performance degrades significantly when labeled device pairs are unavailable. Lastly, they are not robust to uncertainties in linking browsing histories to devices. Binghui Wang, Song Li 0006, Yinzhi Cao, Neil Zhenqiang Gong |
AsiaCCS | 1 |
| 2022 | Cross-domain Prototype Learning from Contaminated Faces via Disentangling Latent FactorsabstractThis paper focuses on an emerging challenging problem called heterogeneous prototype learning (HPL) across face domains-It aims to learn the variation-free target domain prototype for a contaminated input image from the source domain and meanwhile preserve the personal identity. HPL involves two coupled subproblems, i.e., domain transfer and prototype learning. To address the two subproblems in a unified manner, we advocate disentangling the prototype and domain factors in their respected latent feature spaces, and replace the latent source domain features with the target domain ones to generate the heterogeneous prototype. To this end, we propose a disentangled heterogeneous prototype learning framework, dubbed DisHPL, which consists of one encoder-decoder generator and two discriminators. The generator and discriminators play adversarial games such that the generator learns to embed the contaminated image into a prototype feature space only capturing identity information and a domain-specific feature space, as well as generating a realistic-looking heterogeneous prototype. The two discriminators aim to predict personal identities and distinguish between real prototypes versus fake generated prototypes in the source/target domain. Experiments on various heterogeneous face datasets validate the effectiveness of DisHPL. Binghui Wang, Shengbo Chen, Yiu-Ming Cheung, Wei Huang 0013 |
CIKM | 2 |
| 2022 | Bandits for Structure Perturbation-based Black-box Attacks to Graph Neural Networks with Theoretical GuaranteesabstractGraph neural networks (GNNs) have achieved state-of-the-art performance in many graph-based tasks such as node classification and graph classification. However, many recent works have demonstrated that an attacker can mislead GNN models by slightly perturbing the graph structure. Existing attacks to GNNs are either under the less practical threat model where the attacker is assumed to access the GNN model parameters, or under the practical black-box threat model but consider perturbing node features that are shown to be not enough effective. In this paper, we aim to bridge this gap and consider black-box attacks to GNNs with structure perturbation as well as with theoretical guarantees. We propose to address this challenge through bandit techniques. Specifically, we formulate our attack as an online optimization with bandit feedback. This original problem is essentially NP-hard due to the fact that perturbing the graph structure is a binary optimization problem. We then propose an online attack based on bandit optimization which is proven to be sublinear to the query number T, i.e., O(✓NT3/4) where N is the number of nodes in the graph. Finally, we evaluate our proposed attack by conducting experiments over multiple datasets and GNN models. The experimental results on various citation graphs and image graphs show that our attack is both effective and efficient. Binghui Wang, Youqi Li, Pan Zhou 0001 |
CVPR | 1 |
| 2022 | UniCR: Universally Approximated Certified Robustness via Randomized Smoothing
Hanbin Hong, Binghui Wang, Yuan Hong 0001 |
ECCV (5) | 2 |
| 2022 | GraphFL: A Federated Learning Framework for Semi-Supervised Node Classification on GraphsabstractGraph-based semi-supervised node classification (GraphSSC) has wide applications, ranging from networking and security to data mining and machine learning, etc. However, existing centralized GraphSSC methods are impractical to solve many real-world graph-based problems, as collecting the entire graph and labeling a reasonable number of labels is time-consuming and costly, and data privacy may be also violated. Federated learning (FL) is an emerging learning paradigm that enables collaborative learning among multiple clients, which can mitigate the issue of label scarcity and protect data privacy as well. Therefore, performing GraphSSC under the FL setting is a promising solution to solve real-world graph-based problems. However, existing FL methods 1) perform poorly when data across clients are non independent identically distributed (nonIID), 2) cannot handle data with new label domains, and 3) cannot leverage unlabeled data, while all these issues naturally happen in real-world graph-based problems. To address the above issues, we propose the first FL framework, namely GraphFL, for semi-supervised node classification on graphs. Our framework is motivated by meta-learning methods. Specifically, we propose two GraphFL methods to respectively address the non-IID issue in graph data and handle the tasks with new label domains. Furthermore, we design a self-training method to leverage unlabeled graph data. We adopt representative graph neural networks as GraphSSC methods and evaluate GraphFL on multiple graph datasets. Experimental results on various benchmark datasets demonstrate that GraphFL significantly outperforms the compared FL baseline, GraphFL can handle data with new label domains, and GraphFL with selftraining can obtain better performance. Source code is available at https://github.com/binghuiivang/GraphFL. Binghui Wang, Ang Li 0005, Hai Li 0001, Yiran Chen 0001 |
ICDM | 1 |
| 2022 | Almost Tight L0-norm Certified Robustness of Top-k Predictions against Adversarial Perturbations
Jinyuan Jia 0001, Binghui Wang, Hongbin Liu 0005, Neil Zhenqiang Gong |
ICLR | 2 |
| 2022 | Variance of the Gradient Also Matters: Privacy Leakage from GradientsabstractDistributed machine learning (DML) enables model training on a large corpus of decentralized data from users and only collects local models or gradients for global synchronization on the cloud. Recent studies show that a third party can recover the training data in the DML system through publicly shared gradients. Our investigation has revealed that existing techniques (e.g., DLG) can only recover the training data on uniform weight distribution and fail to recover the training data on other weights initialization (e.g., normal distribution) or during the training stage. In this work, we provide an analysis of how weight distribution can affect the training data recovery from gradients. Based on this analysis, we propose a self-adaptive privacy attack from gradients, SAPAG—a general gradient attack algorithm that can recover the training data in DML with any weight initialization and in any training phase. Our algorithm exploits not only the gradients but also the variance of gradients. Specifically, we exploit the variance of gradients distribution and the Deep Neural Network (DNN) architecture and design an adaptive Gaussian kernel of gradient difference as a distance measure. Our experimental results on various benchmark datasets and tasks demonstrate the generalizability of SAPAG. SAPAG outperforms the state-of-the-art algorithms in terms of both the data recovery performance and the recovery speed. Yijue Wang, Jieren Deng, Chenghong Wang, Xianrui Meng, Hang Liu 0001, Binghui Wang, Qin Cao, Caiwen Ding, Sanguthevar Rajasekaran |
IJCNN | 8 |
| 2022 | A Unified Framework for Bidirectional Prototype Learning From Contaminated Faces Across Heterogeneous DomainsabstractExisting heterogeneous face synthesis (HFS) methods focus on performing accurate image-to-image translation across domains, while they cannot effectively remove the nuisance facial variations such as poses, expressions or occlusions. To address such challenges, this paper studies a new practical heterogeneous prototype learning (HPL) problem. To be specific, given a face image contaminated by facial variations from a source domain, HPL aims to reconstruct the variation-free prototype in a specified target domain. To tackle HPL, we propose a unified and end-to-end framework named bidirectional heterogeneous prototype learning (BHPL). As a bidirectional learning framework, BHPL is able to simultaneously reconstruct the heterogeneous prototypes acrosssource-to-targetas well astarget-to-sourcedomains. Furthermore, BHPL is capable of learning the identity prototype features for the contaminated face images from both source and target domains in order to perform robust heterogeneous face recognition. BHPL consists of an encoder-decoder structural generator and two dual-task discriminators, which play an adversarial game such that the generator learns the identity prototype feature and generates the cross-domain identity-preserved prototype for each input face image from both domains, and the discriminators accurately predict face identity and distinguish real versus fake prototypes. Empirically studies on multiple heterogeneous face datasets containing facial variations demonstrate the effectiveness of BHPL. Binghui Wang, Siyu Huang, Yiu-Ming Cheung, Bihan Wen |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2021 | Semi-Supervised Node Classification on Graphs: Markov Random Fields vs. Graph Neural NetworksabstractSemi-supervised node classification on graph-structured data has many applications such as fraud detection, fake account and review detection, user’s private attribute inference in social networks, and community detection. Various methods such as pairwise Markov Random Fields (pMRF) and graph neural networks were developed for semi-supervised node classification. pMRF is more efficient than graph neural networks. However, existing pMRF-based methods are less accurate than graph neural networks, due to a key limitation that they assume a heuristics-based constant edge potential for all edges. In this work, we aim to address the key limitation of existing pMRF-based methods. In particular, we propose to learn edge potentials for pMRF. Our evaluation results on various types of graph datasets show that our optimized pMRF-based method consistently outperforms existing graph neural networks in terms of both accuracy and efficiency. Our results highlight that previous work may have underestimated the power of pMRF for semi-supervised node classification. Binghui Wang, Jinyuan Jia 0001, Neil Zhenqiang Gong |
AAAI | 1 |
| 2021 | On Detecting Growing-Up Behaviors of Malicious Accounts in Privacy-Centric Mobile Social NetworksabstractPrivacy-centric mobile social network (PC-MSN), which allows users to build intimate and private social circles, is an increasingly popular type of online social networks (OSNs). Because of strict usage policy enforced by PC-MSNs (such as restricted account and content access), malicious accounts (or users) have to act like normal accounts to accumulate credentials before committing malicious activities. Therefore, analysis merely relying on static account profile information or social graphs is ineffective to detect such growing-up accounts. Besides, existing behavior-based malicious account detection methods fail to effectively detect growing-up accounts who pretend to be benign and have similar behaviors to benign users during the growing-up stage. Zijie Yang, Binghui Wang, Dong Yuan 0006, Zhuotao Liu, Neil Zhenqiang Gong, Chang Liu 0021, Qi Li 0002, Shaofeng Hu |
ACSAC | 2 |
| 2021 | Robust and Verifiable Information Embedding Attacks to Deep Neural Networks via Error-Correcting CodesabstractIn the era of deep learning, a user often leverages a third-party machine learning tool to train a deep neural network (DNN) classifier and then deploys the classifier as an end-user software product (e.g., a mobile app) or a cloud service. In an information embedding attack, an attacker is the provider of a malicious third-party machine learning tool. The attacker embeds a message into the DNN classifier during training and recovers the message via querying the API of the black-box classifier after the user deploys it. Information embedding attacks have attracted growing attention because of various applications such as watermarking DNN classifiers and compromising user privacy. State-of-the-art information embedding attacks have two key limitations: 1) they cannot verify the correctness of the recovered message, and 2) they are not robust against post-processing (e.g., compression) of the classifier. Jinyuan Jia 0001, Binghui Wang, Neil Zhenqiang Gong |
AsiaCCS | 2 |
| 2021 | A Hard Label Black-box Adversarial Attack Against Graph Neural NetworksabstractGraph Neural Networks (GNNs) have achieved state-of-the-art performance in various graph structure related tasks such as node classification and graph classification. However, GNNs are vulnerable to adversarial attacks. Existing works mainly focus on attacking GNNs for node classification; nevertheless, the attacks against GNNs for graph classification have not been well explored. Jiaming Mu, Binghui Wang, Qi Li 0002, Kun Sun 0001, Mingwei Xu 0001, Zhuotao Liu |
CCS | 2 |
| 2021 | Soteria: Provable Defense Against Privacy Leakage in Federated Learning From Representation PerspectiveabstractFederated learning (FL) is a popular distributed learning framework that can reduce privacy risks by not explicitly sharing private data. However, recent works have demonstrated that sharing model updates makes FL vulnerable to inference attack. In this work, we show our key observation that the data representation leakage from gradients is the essential cause of privacy leakage in FL. We also provide an analysis of this observation to explain how the data presentation is leaked. Based on this observation, we propose a defense called Soteria against model inversion attack in FL. The key idea of our defense is learning to perturb data representation such that the quality of the reconstructed data is severely degraded, while FL performance is maintained. In addition, we derive a certified robustness guarantee to FL and a convergence guarantee to FedAvg, after applying our defense. To evaluate our defense, we conduct experiments on MNIST and CIFAR10 for defending against the DLG attack and GS attack. Without sacrificing accuracy, the results demonstrate that our proposed defense can increase the mean squared error between the reconstructed data and the raw data by as much as 160× for both DLG attack and GS attack, compared with baseline defense methods. Therefore, the privacy of the FL system is significantly improved. Our code can be found at https://github.com/jeremy313/Soteria. Jingwei Sun 0002, Ang Li 0005, Binghui Wang, Huanrui Yang, Hai Li 0001, Yiran Chen 0001 |
CVPR | 3 |
| 2021 | Disentangling Prototype and Variation for Single Sample Face RecognitionabstractSingle sample per person face recognition (SSPP FR) is one of the most challenging problems in FR due to the extreme lack of enrolment data. State-of-the-art SSPP FR methods are based on the prototype plus variation (i.e., P+V) model. However, the classic P+V model has two major limitations: 1) It is a linear model and cannot generalize many non-linear variations; 2) It can be severely impaired once the enrolment face images are contaminated with variations. To this end, we propose a novel disentangled prototype plus variation model, dubbed DisP+V, to tackle such limitations. DisP+V consists of an encoder-decoder structural generator and two discriminators. The generator and discriminators play two adversarial games such that the generator nonlinearly encodes the images into a latent semantic space, where the more discriminative prototype feature and the less discriminative variation feature are disentangled. Meanwhile, the prototype and variation features in the latent space can guide the generator to generate an identity-preserved prototype and the corresponding variation, respectively. Experiments on various real-world face datasets demonstrate the superiority of our DisP+V model over the classic P+V model for SSPP FR. Furthermore, DisP+V demonstrates its unique characteristics in the challenging prototype recovery task. Binghui Wang, Mang Ye, Yiran Chen 0001, Bihan Wen |
ICME | 2 |
| 2021 | LotteryFL: Empower Edge Intelligence with Personalized and Communication-Efficient Federated Learning
Ang Li 0005, Jingwei Sun 0002, Binghui Wang, Lin Duan, Sicheng Li 0001, Yiran Chen 0001, Hai Li 0001 |
SEC | 3 |
| 2021 | A 3-6GHz 5-to-512 Multiplier Adaptive Fast-Locking Self-Biased PLL in 28nm CMOSabstractThis paper presents a design approach for fast- locking and low jitter self-biased phase-locked loop (PLL). The charge-pump current injection technology with minimum area overhead is adopted to accelerate the loop equilibrium capture process without sacrificing the jitter performance. A start-up circuit is proposed in order to shorten the initial ramping up interval of the voltage-controlled oscillator (VCO), which will also help in reducing the lock-in time. A proportional coefficient is introduced in designing self-biased PLLs, which provides more flexibility regarding practical circuits design. The proposed fast- locking self-biased PLL is designed in a TSMC 28nm CMOS process with a supply voltage of 0.9 V. The simulation results show that the locking time is reduced by up to 85% for large division ratios and should not deteriorate the capture performance in small division ratios. Meanwhile, the system almost has no increase in area and the locking time is reduced from 24us to about 3us when operating at 6GHz. Binghui Wang, Haigang Yang, Yiping Jia |
ISCAS | 1 |
| 2021 | Unveiling Fake Accounts at the Time of Registration: An Unsupervised ApproachabstractOnline social networks (OSNs) are plagued by fake accounts. Existing fake account detection methods either require a manually labeled training set, which is time-consuming and costly, or rely on rich information of OSN accounts, e.g., content and behaviors, which incurs significant delay in detecting fake accounts. In this work, we propose UFA (Unveiling Fake Accounts) to detect fake accounts immediately after they are registered in an unsupervised fashion. First, through a measurement study on the registration patterns on a real-world registration dataset, we observe that fake accounts tend to cluster on outlier registration patterns, e.g., IP and phone numbers. Then, we design an unsupervised learning algorithm to learn weights for all registration accounts and their features that reveal outlier registration patterns. Next, we construct a registration graph to capture the correlation between registration accounts, and utilize a community detection method to detect fake accounts via analyzing the registration graph structure. We evaluate UFA using real-world WeChat datasets. Our results demonstrate that UFA achieves a precision 94% with a recall ~80%, while a supervised variant requires 600K manual labels to obtain the comparable performance. Moreover, UFA has been deployed by WeChat to detect fake accounts for more than one year. UFA detects 500K fake accounts per day with a precision ~93% on average, via manual verification by the WeChat security team. Binghui Wang, Shaofeng Hu, Zijie Yang, Dong Yuan 0006, Neil Zhenqiang Gong, Qi Li 0002 |
KDD | 3 |
| 2021 | Privacy-Preserving Representation Learning on Graphs: A Mutual Information PerspectiveabstractLearning with graphs has attracted significant attention recently. Existing representation learning methods on graphs have achieved state-of-the-art performance on various graph-related tasks such as node classification, link prediction, etc. However, we observe that these methods could leak serious private information. For instance, one can accurately infer the links (or node identity) in a graph from a node classifier (or link predictor) trained on the learnt node representations by existing methods. To address the issue, we propose a privacy-preserving representation learning framework on graphs from the mutual information perspective. Specifically, our framework includes a primary learning task and a privacy protection task, and we consider node classification and link prediction as the two tasks of interest. Our goal is to learn node representations such that they can be used to achieve high performance for the primary learning task, while obtaining performance for the privacy protection task close to random guessing. We formally formulate our goal via mutual information objectives. However, it is intractable to compute mutual information in practice. Then, we derive tractable variational bounds for the mutual information terms, where each bound can be parameterized via a neural network. Next, we train these parameterized neural networks to approximate the true mutual information and learn privacy-preserving node representations. We finally evaluate our framework on various graph datasets. Binghui Wang, Ang Li 0005, Yiran Chen 0001, Hai Li 0001 |
KDD | 1 |
| 2021 | Certified Robustness of Graph Neural Networks against Adversarial Structural PerturbationabstractGraph neural networks (GNNs) have recently gained much attention for node and graph classification tasks on graph-structured data. However, multiple recent works showed that an attacker can easily make GNNs predict incorrectly via perturbing the graph structure, i.e., adding or deleting edges in the graph. We aim to defend against such attacks via developing certifiably robust GNNs. Specifically, we prove the first certified robustness guarantee of any GNN for both node and graph classifications against structural perturbation. Moreover, we show that our certified robustness guarantee is tight. Our results are based on a recently proposed technique called randomized smoothing, which we extend to graph data. We also empirically evaluate our method for both node and graph classifications on multiple GNNs and multiple benchmark datasets. For instance, on the Cora dataset, Graph Convolutional Network with our randomized smoothing can achieve a certified accuracy of 0.49 when the attacker can arbitrarily add/delete at most 15 edges in the graph. Binghui Wang, Jinyuan Jia 0001, Neil Zhenqiang Gong |
KDD | 1 |
| 2021 | Towards Adversarial Patch Analysis and Certified Defense against Crowd CountingabstractCrowd counting has drawn much attention due to its importance in safety-critical surveillance systems. Especially, deep neural network (DNN) methods have significantly reduced estimation errors for crowd counting missions. Recent studies have demonstrated that DNNs are vulnerable to adversarial attacks, i.e., normal images with human-imperceptible perturbations could mislead DNNs to make false predictions. In this work, we propose a robust attack strategy called Adversarial Patch Attack with Momentum (APAM) to systematically evaluate the robustness of crowd counting models, where the attacker's goal is to create an adversarial perturbation that severely degrades their performances, thus leading to public safety accidents (e.g., stampede accidents). Especially, the proposed attack leverages the extreme-density background information of input images to generate robust adversarial patches via a series of transformations (e.g., interpolation, rotation, etc.). We observe that by perturbing less than 6% of image pixels, our attacks severely degrade the performance of crowd counting systems, both digitally and physically. To better enhance the adversarial robustness of crowd counting models, we propose the first regression model-based Randomized Ablation (RA), which is more sufficient than Adversarial Training (ADT) (Mean Absolute Error of RA is 5 lower than ADT on clean samples and 30 lower than ADT on adversarial examples). Extensive experiments on five crowd counting models demonstrate the effectiveness and generality of the proposed method. Zhikang Zou, Pan Zhou 0001, Xiaoqing Ye, Binghui Wang, Ang Li 0005 |
ACM Multimedia | 5 |
| 2021 | Backdoor Attacks to Graph Neural NetworksabstractIn this work, we propose the first backdoor attack to graph neural networks (GNN). Specifically, we propose a subgraph based backdoor attack to GNN for graph classification. In our backdoor attack, a GNN classifier predicts an attacker-chosen target label for a testing graph once a predefined subgraph is injected to the testing graph. Our empirical results on three real-world graph datasets show that our backdoor attacks are effective with a small impact on a GNN's prediction accuracy for clean testing graphs. Moreover, we generalize a randomized smoothing based certified defense to defend against our backdoor attacks. Our empirical results show that the defense is effective in some cases but ineffective in other cases, highlighting the needs of new defenses for our backdoor attacks. Zaixi Zhang, Jinyuan Jia 0001, Binghui Wang, Neil Zhenqiang Gong |
SACMAT | 3 |
| 2021 | VD-GAN: A Unified Framework for Joint Prototype and Representation Learning From Contaminated Single Sample per PersonabstractSingle sample per person (SSPP) face recognition with a contaminated biometric enrolment database (SSPP-ce FR) is an emerging practical FR problem, where the SSPP in the enrolment database is no longer standard but contaminated by nuisance facial variations such as expression, lighting, pose, and disguise. In this case, the conventional SSPP FR methods, including the patch-based and generic learning methods, will suffer from serious performance degradation. Few recent methods were proposed to tackle SSPP-ce FR by either performing prototype learning on the contaminated enrolment database or learning discriminative representations that are robust against variation. Despite that, most of these approaches can only handle a specified single variation, e.g., pose, but cannot be extended to multiple variations. To address these two limitations, we propose a novel Variation Disentangling Generative Adversarial Network (VDGAN) to jointly perform prototype learning and representation learning in a unified framework. The proposed VD-GAN consists of an encoder-decoder structural generator and a multi-task discriminator to handle universal variations including single, multiple, and even mixed variations in practice. The generator and discriminator play an adversarial game such that the generator learns a discriminative identity representation and generates an identity-preserved prototype for each face image, while the discriminator aims to predict face identity label, distinguish real vs. fake prototype, and disentangle target variations from the learned representations. Qualitative and quantitative evaluations on various real-world face datasets containing single/multiple and mixed variations demonstrate the effectiveness of VD-GAN. Binghui Wang, Yiu-Ming Cheung, Yiran Chen 0001, Bihan Wen |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2020 | Certified Robustness for Top-k Predictions against Adversarial Perturbations via Randomized Smoothing
Jinyuan Jia 0001, Binghui Wang, Neil Zhenqiang Gong |
ICLR | 3 |
| 2020 | Perturbing Across the Feature Hierarchy to Improve Standard and Strict Blackbox Attack TransferabilityabstractWe consider the blackbox transfer-based targeted adversarial attack threat model in the realm of deep neural network (DNN) image classifiers. Rather than focusing on crossing decision boundaries at the output layer of the source model, our method perturbs representations throughout the extracted feature hierarchy to resemble other classes. We design a flexible attack framework that allows for multi-layer perturbations and demonstrates state-of-the-art targeted transfer performance between ImageNet DNNs. We also show the superiority of our feature space methods under a relaxation of the common assumption that the source and target models are trained on the same dataset and label space, in some instances achieving a $10\times$ increase in targeted success rate relative to other blackbox transfer methods. Finally, we analyze why the proposed methods outperform existing attack strategies and show an extension of the method in the case when limited queries to the blackbox model are allowed. Nathan Inkawhich, Kevin J. Liang, Binghui Wang, Matthew Inkawhich, Lawrence Carin, Yiran Chen 0001 |
NeurIPS | 3 |
| 2020 | Certified Robustness of Community Detection against Adversarial Structural Perturbation via Randomized SmoothingabstractCommunity detection plays a key role in understanding graph structure. However, several recent studies showed that community detection is vulnerable to adversarial structural perturbation. In particular, via adding or removing a small number of carefully selected edges in a graph, an attacker can manipulate the detected communities. However, to the best of our knowledge, there are no studies on certifying robustness of community detection against such adversarial structural perturbation. In this work, we aim to bridge this gap. Specifically, we develop the first certified robustness guarantee of community detection against adversarial structural perturbation. Given an arbitrary community detection method, we build a new smoothed community detection method via randomly perturbing the graph structure. We theoretically show that the smoothed community detection method provably groups a given arbitrary set of nodes into the same community (or different communities) when the number of edges added/removed by an attacker is bounded. Moreover, we show that our certified robustness is tight. We also empirically evaluate our method on multiple real-world graphs with ground truth communities. Jinyuan Jia 0001, Binghui Wang, Neil Zhenqiang Gong |
WWW | 2 |
| 2020 | Synergistic Generic Learning for Face Recognition From a Contaminated Single Sample per PersonabstractSingle sample per person face recognition (SSPP FR), i.e., identifying a person (i.e., data subject) with a single face image only for training, has several attractive potential applications, but it is still a challenging problem. Existing generic learning methods usually leverage prototype plus variation (P+V) model for SSPP FR provided that face samples in the biometric enrolment database are variation-free and thus can be treated as the prototypes of data subjects. However, this condition is not satisfied when these samples are contaminated by nuisance facial variations in the wild, such as varied expressions, poor lightings, and disguises (e.g., wearing scarf). We call this new and practical problem SSPP FR with a contaminated biometric enrolment database (SSPP-ce FR). Subsequently, a challenging issue will be raised on estimating proper prototypes from the contaminated enrolment samples in SSPP-ce FR. Moreover, the generated variation dictionary also needs to be enhanced because it is simply based on the subtraction of average face from the samples of the same data subject in the generic set, thus containing individual characteristics that can hardly be shared by other data subjects. To address these two issues, we propose a novel synergistic generic learning (SGL) method to study the SSPP-ce FR problem. Compared with the existing generic learning methods, SGL develops a new “learned P + learned V” model to identify new query samples. Specifically, it learns better prototypes for the contaminated samples in the biometric enrolment database by preserving their more discriminative subject-specific portions and learns a representative variation dictionary by extracting the less discriminative intra-subject variants from an auxiliary generic set. The experiments on various benchmark face datasets demonstrate the effectiveness of the proposed SGL method. Yiu-Ming Cheung, Binghui Wang, Jian Lou 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2019 | Attacking Graph-based Classification via Manipulating the Graph StructureabstractGraph-based classification methods are widely used for security analytics. Roughly speaking, graph-based classification methods include collective classification and graph neural network. Attacking a graph-based classification method enables an attacker to evade detection in security analytics. However, existing adversarial machine learning studies mainly focused on machine learning for non-graph data. Only a few recent studies touched adversarial graph-based classification methods. However, they focused on graph neural network, leaving collective classification largely unexplored. We aim to bridge this gap in this work. We consider an attacker's goal is to evade detection via manipulating the graph structure. We formulate our attack as a graph-based optimization problem, solving which produces the edges that an attacker needs to manipulate to achieve its attack goal. However, it is computationally challenging to solve the optimization problem exactly. To address the challenge, we propose several approximation techniques to solve the optimization problem. We evaluate our attacks and compare them with a recent attack designed for graph neural networks using four graph datasets. Our results show that our attacks can effectively evade graph-based classification methods. Moreover, our attacks outperform the existing attack for evading collective classification methods and some graph neural network methods. Binghui Wang, Neil Zhenqiang Gong |
CCS | 1 |
| 2019 | Graph-based Security and Privacy Analytics via Collective Classification with Joint Weight Learning and Propagation
Binghui Wang, Jinyuan Jia 0001, Neil Zhenqiang Gong |
NDSS | 1 |
| 2019 | Robust heterogeneous discriminative analysis for face recognition with single sample per person
Yiu-Ming Cheung, Binghui Wang, Risheng Liu |
Pattern Recognit. | 3 |
| 2018 | SybilBlind: Detecting Fake Users in Online Social Networks Without Manual Labels
Binghui Wang, Le Zhang 0006, Neil Zhenqiang Gong |
RAID | 1 |
| 2018 | Stealing Hyperparameters in Machine LearningabstractHyperparameters are critical in machine learning, as different hyperparameters often result in models with significantly different performance. Hyperparameters may be deemed confidential because of their commercial value and the confidentiality of the proprietary algorithms that the learner uses to learn them. In this work, we propose attacks on stealing the hyperparameters that are learned by a learner. We call our attacks hyperparameter stealing attacks. Our attacks are applicable to a variety of popular machine learning algorithms such as ridge regression, logistic regression, support vector machine, and neural network. We evaluate the effectiveness of our attacks both theoretically and empirically. For instance, we evaluate our attacks on Amazon Machine Learning. Our results demonstrate that our attacks can accurately steal hyperparameters. We also study countermeasures. Our results highlight the need for new defenses against our hyperparameter stealing attacks for certain machine learning algorithms. Binghui Wang, Neil Zhenqiang Gong |
IEEE Symposium on Security and Privacy | 1 |
| 2017 | Robust Heterogeneous Discriminative Analysis for Single Sample Per Person Face RecognitionabstractSingle sample face recognition is one of the most challenging problems in face recognition (FR), where only one single sample per person (SSPP) is enrolled in the gallery set for training. Although patch-based methods have achieved great success in FR with SSPP, they still have significant limitations. In this work, we propose a new patch-based method, namely Robust Heterogeneous Discriminative Analysis (RHDA), to tackle FR with SSPP. Compared with the existing patch-based methods, RHDA can enhance the robustness against complex facial variations from two aspects. First, we develop a novel Fisher-like criterion, which incorporates two manifold embeddings, to learn heterogeneous discriminative representations of image patches. Specifically, for each patch, the Fisher-like criterion is able to preserve the reconstruction relationship of neighboring patches from the same person, while suppressing neighboring patches from different persons. Second, we present two distance metrics, i.e., patch-to-patch distance and patch-to-manifold distance, and develop a fusion strategy to combine the recognition outputs of above two distance metrics via joint majority voting for identification. Experimental results on the AR and FERET benchmark datasets demonstrate the efficacy of the proposed method. Yiu-Ming Cheung, Binghui Wang, Risheng Liu |
CIKM | 3 |
| 2017 | Random Walk Based Fake Account Detection in Online Social NetworksabstractOnline social networks are known to be vulnerable to the so-called Sybil attack, in which an attacker maintains massive fake accounts (also called Sybils) and uses them to perform various malicious activities. Therefore, Sybil detection is a fundamental security research problem in online social networks. Random walk based methods, which leverage the structure of an online social network to distribute reputation scores for users, have been demonstrated to be promising in certain real-world online social networks. In particular, random walk based methods have three desired features: they can have theoretically guaranteed performance for online social networks that have the fast-mixing property, they are accurate when the social network has strong homophily property, and they can be scalable to large-scale online social networks. However, existing random walk based methods suffer from several key limitations: 1) they can only leverage either labeled benign users or labeled Sybils, but not both, 2) they have limited detection accuracy for weak-homophily social networks, and 3) they are not robust to label noise in the training dataset. In this work, we propose a new random walk based Sybil detection method called SybilWalk. SybilWalk addresses the limitations of existing random walk based methods while maintaining their desired features. We perform both theoretical and empirical evaluations to compare SybilWalk with previous random walk based methods. Theoretically, for online social networks with the fast-mixing property, SybilWalk has a tighter asymptotical bound on the number of Sybils that are falsely accepted into the social network than all existing random walk based methods. Empirically, we compare SybilWalk with previous random walk based methods using both social networks with synthesized Sybils and a large-scale Twitter dataset with real Sybils. Our empirical results demonstrate that 1) SybilWalk is substantially more accurate than existing random walk based methods for weakhomophily social networks, 2) SybilWalk is substantially more robust to label noise than existing random walk based methods, and 3) SybilWalk is as scalable as the most efficient existing random walk based methods. In particular, on the Twitter dataset, SybilWalk achieves a false positive rate of 1.3% and a false negative rate of 17.3%. Jinyuan Jia 0001, Binghui Wang, Neil Zhenqiang Gong |
DSN | 2 |
| 2017 | GANG: Detecting Fraudulent Users in Online Social Networks via Guilt-by-Association on Directed GraphsabstractDetecting fraudulent users in online social networks is a fundamental and urgent research problem as adversaries can use them to perform various malicious activities. Global social structure based methods, which are known as guilt-by-association, have been shown to be promising at detecting fraudulent users. However, existing guilt-by-association methods either assume symmetric (i.e., undirected) social links, which oversimplifies the asymmetric (i.e., directed) social structure of real-world online social networks, or only leverage labeled fraudulent users or labeled normal users (but not both) in the training dataset, which limits detection accuracies. In this work, we propose GANG, a guilt-by-association method on directed graphs, to detect fraudulent users in OSNs. GANG is based on a novel pairwise Markov Random Field that we design to capture the unique characteristics of the fraudulent-user-detection problem in directed OSNs. In the basic version of GANG, given a training dataset, we leverage Loopy Belief Propagation (LBP) to estimate the posterior probability distribution for each user and uses it to predict a user's label. However, the basic version is not scalable enough and not guaranteed to converge because it relies on LBP. Therefore, we further optimize GANG and our optimized version can be represented as a concise matrix form, with which we are able to derive conditions for convergence. We compare GANG with various existing guilt-by-association methods on a large-scale Twitter dataset and a large-scale Sina Weibo dataset with labeled fraudulent and normal users. Our results demonstrate that GANG substantially outperforms existing methods, and that the optimized version of GANG is significantly more efficient than the basic version. Binghui Wang, Neil Zhenqiang Gong, Hao Fu 0015 |
ICDM | 1 |
| 2017 | SybilSCAR: Sybil detection in online social networks via local rule based propagationabstractDetecting Sybils in online social networks (OSNs) is a fundamental security research problem as adversaries can leverage Sybils to perform various malicious activities. Structure-based methods have been shown to be promising at detecting Sybils. Existing structure-based methods can be classified into two categories: Random Walk (RW)-based methods and Loop Belief Propagation (LBP)-based methods. RW-based methods cannot leverage labeled Sybils and labeled benign users simultaneously, which limits their detection accuracy, and they are not robust to noisy labels. LBP-based methods are not scalable, and they cannot guarantee convergence. In this work, we propose SybilSCAR, a new structure-based method to perform Sybil detection in OSNs. SybilSCAR maintains the advantages of existing methods while overcoming their limitations. Specifically, SybilSCAR is Scalable, Convergent, Accurate, and Robust to label noises. We first propose a framework to unify RW-based and LBP-based methods. Under our framework, these methods can be viewed as iteratively applying a (different) local rule to every user, which propagates label information among a social graph. Second, we design a new local rule, which SybilSCAR iteratively applies to every user to detect Sybils. We compare SybilSCAR with a state-of-the-art RW-based method and a state-of-the-art LBP-based method, using both synthetic Sybils and large-scale social network datasets with real Sybils. Our results demonstrate that SybilSCAR is more accurate and more robust to label noise than the compared state-of-the-art RW-based method, and that SybilSCAR is orders of magnitude more scalable than the state-of-the-art LBP-based method and is guaranteed to converge. To facilitate research on Sybil detection, we have made our implementation of SybilSCAR publicly available on our webpages. Binghui Wang, Le Zhang 0006, Neil Zhenqiang Gong |
INFOCOM | 1 |
| 2017 | AttriInfer: Inferring User Attributes in Online Social Networks Using Markov Random FieldsabstractIn the attribute inference problem, we aim to infer users' private attributes (e.g., locations, sexual orientation, and interests) using their public data in online social networks. State-of-the-art methods leverage a user's both public friends and public behaviors (e.g., page likes on Facebook, apps that the user reviewed on Google Play) to infer the user's private attributes. However, these methods suffer from two key limitations: 1) suppose we aim to infer a certain attribute for a target user using a training dataset, they only leverage the labeled users who have the attribute, while ignoring the label information of users who do not have the attribute; 2) they are inefficient because they infer attributes for target users one by one. As a result, they have limited accuracies and applicability in real-world social networks. Jinyuan Jia 0001, Binghui Wang, Le Zhang 0006, Neil Zhenqiang Gong |
WWW | 2 |
| 2016 | Discriminant Manifold Learning via Sparse Coding for Image Analysis
Binghui Wang, Xin Fan 0001, Chuang Lin 0001 |
MMM (2) | 2 |
| 2014 | Neighbourhood sensitive preserving embedding for pattern classificationabstractRecently, a large family of supervised or unsupervised manifold learning algorithms that stem from statistical or geometrical theory has been designed to solve the problem of pattern classification. In this study, consider the fact that the data are usually sampled from a low‐dimensional manifold space which resides in a high‐dimensional Euclidean space, the authors propose a novel two‐graph‐based supervised linear classification algorithm called neighbourhood sensitive preserving embedding (NSPE). Different from local linear embedding (LLE) (or neighbourhood preserving embedding (NPE)) which preserves the local neighbourhood structure with one graph, NSPE can discover both the intrinsic and discriminant structure of the data manifold by constructing two graphs, that is, the within‐class graph and the between‐class graph. Thus, the data are mapped into a subspace where the nearby points with the same label are close to each other, whereas the nearby points with different labels are far apart. As a classification method, besides being defined on training samples, NSPE is also defined on testing samples. Experiments carried on the real‐world face databases demonstrate that the results of all two‐graph‐based spectral methods are comparable and better than that of one‐graph‐based methods. Binghui Wang, Chuang Lin 0001, Xue-Feng Zhao, Zheming Lu 0001 |
IET Image Process. | 1 |
| 2014 | Hierarchical Bayes based Adaptive Sparsity in Gaussian Mixture Model
Binghui Wang, Chuang Lin 0001, Xin Fan 0001, Ning Jiang 0001, Dario Farina |
Pattern Recognit. Lett. | 1 |