EDBT 2026 Demo / reviewers in the wild / expert
Haiqin Weng
dblp:169/7167
· DBLP profile ↗
20ranked-venue papers
5as first author
15since 2021 · last 2025
0000-0002-3005-761XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 9 · 9 since 2021Artificial intelligence and machine learning · 6 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Anti-FT: Towards Practical Deep Leakage From GradientsabstractFederated learning is usually regarded as a privacy-preserving training paradigm for it enables multiple clients to participate in a training task without sharing their private data. However, recent studies revealed that a malicious server can still recover private data from the victim clients based on the shared gradients via deep leakage from gradients (DLG). Currently, almost all DLG attacks are designed based on the average loss, leading to a significant decrease in attack efficiency when the batch size is greater than 1. In this paper, we revisit DLG attacks from the perspective of the loss function. We reveal that not all samples in the target batch are equally susceptible to DLG attacks: the sample with the highest loss value tends to be easily recovered by DLG attacks. Based on these observations, we propose a simple yet effective DLG method under practical FL settings. Specifically, the adversaries can enhance the effectiveness of DLG by perturbing the global model through finetuning it with a few mislabeled samples (dubbed ‘Anti-FT’). Extensive experiments are conducted on benchmark datasets, which verify the effectiveness of our method and its resistance to potential defenses. The codes are available at https://github.com/zlh-thu/anti-finetune. Linghui Zhu, Yiming Li 0004, Haiqin Weng, Shutao Xia, Zhi Wang 0001 |
ICIP | 3 |
| 2025 | A Benchmark for Semantic Sensitive Information in LLMs OutputsabstractLarge language models (LLMs) can output sensitive information, which has emerged as a novel safety concern. Previous works focus on structured sensitive information (e.g. personal identifiable information).
However, we notice that sensitive information can also be at semantic level, i.e. semantic sensitive information (SemSI).
Particularly, *simple natural questions* can let state-of-the-art (SOTA) LLMs output SemSI.
%which is hard to be detected compared with structured ones.
Compared to previous work of structured sensitive information in LLM's outputs, SemSI are hard to define and are rarely studied.
Therefore, we propose a novel and large-scale investigation on the existence of SemSI in SOTA LLMs induced by simple natural questions.
First, we construct a comprehensive and labeled dataset of semantic sensitive information, SemSI-Set, by including three typical categories of SemSI.
Then, we propose a large-scale benchmark, SemSI-Bench, to systematically evaluate semantic sensitive information in 25 SOTA LLMs.
Our finding reveals that SemSI widely exists in SOTA LLMs' outputs by querying with simple natural questions.
We open-source our project at https://semsi-project.github.io/. Han Qiu 0001, Yiming Li 0004, Tianwei Zhang 0004, Wenyu Zhu, Haiqin Weng, Liu Yan, Chao Zhang 0008 |
ICLR | 7 |
| 2025 | PatchSegDet: Attack-Agnostic Detection of Physical Adversarial Patches in Face Recognition SystemsabstractAdversarial patch attacks are an emerging security threat for real-world Face Recognition Systems (FRS). Although many adversarial patch detection methods have been proposed for image classification, to the best of our knowledge, few have yet been specifically developed for FRS. Furthermore, the characteristics of FRS attack vectors impede current detection methods from being adapted to FRS. To bridge this gap, we propose PatchSegDet, an attack-agnostic two-stage adversarial patch detection method to safeguard FRS. It employs the Segment Anything Model (SAM) to segment out suspicious features and determines whether they constitute attacks against FRS. Leveraging SAM’s remarkable generalization and zero-shot capabilities in facial image segmentation, PatchSegDet is capable of detecting patches with varying textures and patterns placed in any facial region. Extensive experiments demonstrate the effectiveness and robustness of PatchSegDet against various attack methods in both digital and physical domains. Our findings provide insights for practitioners to better defend physical adversarial patch attacks in real-world FRS. Qinfeng Li, Xuhong Zhang 0002, Xiaochu Chen, Haiqin Weng, Yan Liu 0069 |
ICME | 7 |
| 2025 | RACONTEUR: A Knowledgeable, Insightful, and Portable LLM-Powered Shell Command Explainer
Jiangyi Deng, Xinfeng Li, Yanjiao Chen, Yijie Bai, Haiqin Weng, Yan Liu 0069, Tao Wei 0002, Wenyuan Xu 0001 |
NDSS | 5 |
| 2025 | FDINet: Protecting Against DNN Model Extraction Using Feature Distortion IndexabstractMachine Learning as a Service (MLaaS) platforms have gained popularity due to their accessibility, cost-efficiency, scalability, and rapid development capabilities. However, recent research has highlighted the vulnerability of cloud-based models in MLaaS to model extraction attacks. In this paper, we introduce FDINet, a novel defense mechanism that leverages the feature distribution of deep neural network (DNN) models. Concretely, by analyzing the feature distribution from the adversary's queries, we reveal that the feature distribution of these queries deviates from that of the model's problem domain. Based on this key observation, we propose Feature Distortion Index (FDI), a metric designed to quantitatively measure the feature distribution deviation of received queries. The proposed FDINet utilizes FDI to train a binary detector and exploits FDI similarity to identify colluding adversaries from distributed extraction attacks. We conduct extensive experiments to evaluate FDINet against six state-of-the-art extraction attacks on four benchmark datasets and four popular model architectures. Empirical results demonstrate the following findings: (1) FDINet proves to be highly effective in detecting model extraction, achieving a100% detection accuracyon DFME and DaST. (2) FDINet is highly efficient, using just 50 queries to raise an extraction alarm with anaverage confidence of 96.08%for GTSRB. (3) FDINet exhibits the capability to identify colluding adversaries with an accuracyexceeding 91%. Additionally, it demonstrates the ability to detect two types of adaptive attacks. Hongwei Yao, Zheng Li 0023, Haiqin Weng, Zhan Qin, Kui Ren 0001 |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2024 | ProFake: Detecting Deepfakes in the Wild against Quality Degradation with Progressive Quality-adaptive LearningabstractDespite the promising advances in deepfake detection on current datasets, detecting visual deepfakes in real-world scenarios (e.g., deepfake videos and live streaming on YouTube) remains a challenge due to the inherent quality degradation such as unpredictable compression employed by social media platforms. Such degradation perturbs discernible forgery clues and diminishes the effectiveness of deepfake detection methods, raising a critical safety concern to the misuse of forgery faces in real-world scenarios. In this paper, we aim to understand the impacts of real-world degradation on the robustness of deepfake detection. Particularly, we investigate the risk of degraded deepfakes towards their detection on two real-world scenarios (i.e., deepfake videos and deepfake live streaming on social media platforms). By measuring the effects of real-world degradations on the performance and representation capabilities of detection models, we reveal that real-world deepfakes can be simulated via common degradation operations (e.g., JPEG compression) as they are perceptually similar to deepfake detectors. By analyzing the training dynamics under different sequences of training samples, we observe that the training order of deepfakes progressing from non-degraded (easy) to heavily degraded (hard) enhances the adaptability of detection models to various degradation in real-world scenarios. Drawing from these observations, we present a novel deepfake detection method ProFake to enhance the robustness of deepfake detection against real-world quality degradations. ProFake enables quality-adaptive learning via progressively degrade, detect and assign weights for the training samples driven by the feedback of model performance and image quality, which ensures that our model gradually focuses on more challenging samples to achieve quality-adaptive deepfake detection. Extensive experiments show that compared with existing methods, ProFake improves deepfake detection accuracy by an average of over 10 % in real-world scenarios and by an average of over 30 % in heavily degraded scenarios, while maintaining comparable performance in detecting high-quality deepfakes. Huiyu Xu, Yaopeng Wang, Zhibo Wang 0001, Zhongjie Ba, Haiqin Weng, Tao Wei 0002, Kui Ren 0001 |
CCS | 7 |
| 2024 | Sophon: Non-Fine-Tunable Learning to Restrain Task Transferability For Pre-trained ModelsabstractInstead of building deep learning models from scratch, developers are more and more relying on adapting pre-trained models to their customized tasks. However, powerful pre-trained models may be misused for unethical or illegal tasks, e.g., privacy inference and unsafe content generation. In this paper, we introduce a pioneering learning paradigm, non-fine-tunable learning, which prevents the pre-trained model from being fine-tuned to indecent tasks while preserving its performance on the original task. To fulfill this goal, we propose Sophon, a protection framework that reinforces a given pre-trained model to be resistant to being fine-tuned in pre-defined restricted domains. Nonetheless, this is challenging due to a diversity of complicated fine-tuning strategies that may be adopted by adversaries. Inspired by model-agnostic meta-learning, we overcome this difficulty by designing sophisticated fine-tuning simulation and fine-tuning evaluation algorithms. In addition, we carefully design the optimization process to entrap the pre-trained model within a hard-to-escape local optimum regarding restricted domains. We have conducted extensive experiments on two deep learning modes (classification and generation), seven restricted domains, and six model architectures to verify the effectiveness of Sophon. Experiment results verify that fine-tuning Sophon-protected models incurs an overhead comparable to or even greater than training from scratch. Furthermore, we confirm the robustness of Sophon to three fine-tuning methods, five optimizers, various learning rates and batch sizes. Sophon may help boost further investigations into safe and responsible AI. Jiangyi Deng, Shengyuan Pang, Yanjiao Chen, Liangming Xia, Yijie Bai, Haiqin Weng, Wenyuan Xu 0001 |
SP | 6 |
| 2024 | Exploring ChatGPT's Capabilities on Vulnerability Management
Peiyu Liu 0003, Lirong Fu, Kangjie Lu, Xuhong Zhang 0002, Wenzhi Chen, Haiqin Weng, Shouling Ji, Wenhai Wang |
USENIX Security Symposium | 8 |
| 2023 | Devil in Disguise: Breaching Graph Neural Networks Privacy through InfiltrationabstractGraph neural networks (GNNs) have been developed to mine useful information from graph data of various applications, e.g., healthcare, fraud detection, and social recommendation. However, GNNs open up new attack surfaces for privacy attacks on graph data. In this paper, we propose Infiltrator, a privacy attack that is able to pry node-level private information based on black-box access to GNNs. Different from existing works that require prior information of the victim node, we explore the possibility of conducting the attack without any information of the victim node. Our idea is to infiltrate the graph with attacker-created nodes to befriend the victim node. More specifically, we design infiltration schemes that enable the adversary to infer the label, neighboring links, and sensitive attributes of a victim node. We evaluate Infiltrator with extensive experiments on three representative GNN models and six real-world datasets. The results demonstrate that Infiltrator can achieve an attack performance of more than 98% in all three attacks, outperforming baseline approaches. We further evaluate the defense resistance of Infiltrator against the graph homophily defender and the differentially private model. Lingshuo Meng, Yijie Bai, Yanjiao Chen, Yutong Hu 0005, Wenyuan Xu 0001, Haiqin Weng |
CCS | 6 |
| 2023 | Counterfactual-based Saliency Map: Towards Visual Contrastive Explanations for Neural NetworksabstractExplaining deep models in a human-understandable way has been explored by many works that mostly explain why an input causes a corresponding prediction (i.e., Why P?). However, seldom they could handle those more complex causal questions like "Why P rather than Q?" and "Why one is P, while another is Q?", which would better help humans understand the behavior of deep models. Considering the insufficient study on such complex causal questions, we make the first attempt to explain different causal questions by contrastive explanations in a unified framework, i.e., Counterfactual Contrastive Explanation (CCE), which visually and intuitively explains the aforementioned questions via a novel positive-negative saliency-based explanation scheme. More specifically, we propose a content-aware counterfactual perturbing algorithm to stimulate contrastive examples, from which a pair of positive and negative saliency maps could be derived to contrastively explain why P (positive class) rather than Q (negative class). Beyond existing works, our counterfactual perturbation meets the principles of validity, sparsity, and data distribution closeness at the same time. In addition, by slightly adjusting the objective of perturbation, our framework can adapt to different causal questions. Extensive experimental evaluation demonstrates the effectiveness and superior performance of the proposed CCE on different benchmark metrics for interpretability, including Sanity Check, Class Deviation Score and Insertion-Deletion tests. A user study is conducted and the results show that user confidence is increasing significantly when presented with CCE compared to standard saliency map baselines. Zhibo Wang 0001, Haiqin Weng, Hengchang Guo, Tao Wei 0002, Kui Ren 0001 |
ICCV | 3 |
| 2023 | VILLAIN: Backdoor Attacks Against Vertical Split Learning
Yijie Bai, Yanjiao Chen, Hanlei Zhang, Wenyuan Xu 0001, Haiqin Weng, Dou Goodman |
USENIX Security Symposium | 5 |
| 2022 | Towards Certifying the Asymmetric Robustness for Neural Networks: Quantification and ApplicationsabstractOne intriguing property of deep neural networks (DNNs) is their vulnerability to adversarial examples – those maliciously crafted inputs that deceive target DNNs. While a plethora of defenses have been proposed to mitigate the threats of adversarial examples, they are often penetrated or circumvented by even stronger attacks. To end the constant arms race between attackers and defenders, significant efforts have been devoted to providing certifiable robustness bounds for DNNs, which ensures that for a given input its vicinity does not admit any adversarial instances. Yet, most prior works focus on the case of symmetric vicinities (e.g., a hyperrectangle centered at a given input), while ignoring the inherent heterogeneity of perturbation direction (e.g., the input is more vulnerable along a particular perturbation direction). To bridge the gap, in this article, we propose the concept ofasymmetric robustnessto account for the inherent heterogeneity of perturbation directions, and presentAmoeba1, an efficient certification framework for asymmetric robustness. Through extensive empirical evaluation on state-of-the-art DNNs and benchmark datasets, we show that compared with its symmetric counterpart, the asymmetric robustness bound of a given input describes its local geometric properties in a more precise manner, which enables use cases including (i) modeling stronger adversarial threats, (ii) interpreting DNN predictions, and makes it a more practical definition of certifiable robustness for security-sensitive domains. Changjiang Li, Shouling Ji, Haiqin Weng, Bo Li 0026, Raheem A. Beyah, Shanqing Guo, Zonghui Wang, Ting Wang 0006 |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2022 | A Large-Scale Empirical Study on the Vulnerability of Deployed IoT DevicesabstractThe Internet of Things (IoT) has become ubiquitous and greatly affected peoples’ daily lives. With the increasing development of IoT devices, the corresponding security issues are becoming more and more challenging. Such a severe security situation raises the following questions that need urgent attention: What are the primary security threats that IoT devices face currently? How do vendors and users deal with these threats? In this article, we aim to answer these critical questions through a large-scale systematic study. Specifically, we perform a ten-month-long empirical study on the vulnerability of 1,362,906 IoT devices varying from six types. The results show sufficient evidence that N-days vulnerability is seriously endangering the IoT devices: 385,060 (28.25 percent) devices suffer from at least one N-days vulnerability. Moreover, 2669 of these vulnerable devices may have been compromised by botnets. We further reveal the massive differences among five popular IoT search engines:Shodan[1],Censys[2], [3],Zoomeye[4],Fofa[5], andNTI[6]. To study whether vendors and users adopt defenses against the threats, we measure the security of MQTT [7] servers, and identify that 12740 (88 percent) MQTT servers have no password protection. Our analysis can serve as an important guideline for investigating the security of IoT devices, as well as advancing the development of a more secure environment for IoT systems. Shouling Ji, Wei-Han Lee, Changting Lin, Haiqin Weng, JingZheng Wu, Pan Zhou 0001, Liming Fang 0001, Raheem A. Beyah |
IEEE Trans. Dependable Secur. Comput. | 5 |
| 2021 | Noise Doesn't Lie: Towards Universal Detection of Deep InpaintingabstractDeep image inpainting aims to restore damaged or missing regions in an image with realistic contents. While having a wide range of applications such as object removal and image recovery, deep inpainting techniques also have the risk of being manipulated for image forgery. A promising countermeasure against such forgeries is deep inpainting detection, which aims to locate the inpainted regions in an image. In this paper, we make the first attempt towards universal detection of deep inpainting, where the detection network can generalize well when detecting different deep inpainting methods. To this end, we first propose a novel data generation approach to generate a universal training dataset, which imitates the noise discrepancies exist in real versus inpainted image contents to train universal detectors. We then design a Noise-Image Cross-fusion Network (NIX-Net) to effectively exploit the discriminative information contained in both the images and their noise patterns. We empirically show, on multiple benchmark datasets, that our approach outperforms existing detection methods by a large margin and generalize well to unseen deep inpainting techniques. Our universal training dataset can also significantly boost the generalizability of existing detection methods. Ang Li 0008, Qiuhong Ke, Xingjun Ma, Haiqin Weng, Zhiyuan Zong, Rui Zhang 0003 |
IJCAI | 4 |
| 2021 | Fast-RCM: Fast Tree-Based Unsupervised Rare-Class MiningabstractRare classes are usually hidden in an imbalanced dataset with the majority of the data examples from major classes. Rare-class mining (RCM) aims at extracting all the data examples belonging to rare classes. Most of the existing approaches for RCM require a certain amount of labeled data examples as input. However, they are ineffective in practice since requesting label information from domain experts is time consuming and human-labor extensive. Thus, we investigate the unsupervised RCM problem, which to the best of our knowledge is the first such attempt. To this end, we propose an efficient algorithm called Fast-RCM for unsupervised RCM, which has an approximately linear time complexity with respect to data size and data dimensionality. Given an unlabeled dataset, Fast-RCM mines out the rare class by first building a rare tree for the input dataset and then extracting data examples of the rare classes based on this rare tree. Compared with the existing approaches which have quadric or even cubic time complexity, Fast-RCM is much faster and can be extended to large-scale datasets. The experimental evaluation on both synthetic and real-world datasets demonstrate that our algorithm can effectively and efficiently extract the rare classes from an unlabeled dataset under the unsupervised settings, and is approximately five times faster than that of the state-of-the-art methods. Haiqin Weng, Shouling Ji, Changchang Liu, Ting Wang 0006, Qinming He, Jianhai Chen |
IEEE Trans. Cybern. | 1 |
| 2020 | De-Health: All Your Online Health Information Are Belong to UsabstractIn this paper, we study the privacy of online health data. We present a novel online health data De-Anonymization (DA) framework, named De-Health. Leveraging two real world online health datasets WebMD and HealthBoards, we validate the DA efficacy of De-Health. We also present a linkage attack framework which can link online health/medical information to real world people. Through a proof-of-concept attack, we link 347 out of 2805 WebMD users to real world people, and find the full names, medical/health information, birthdates, phone numbers, and other sensitive information for most of the re-identified users. This clearly illustrates the fragility of the privacy of those who use online health forums. Shouling Ji, Qinchen Gu, Haiqin Weng, Qianjun Liu, Pan Zhou 0001, Jing Chen 0003, Zhao Li 0007, Raheem A. Beyah, Ting Wang 0006 |
ICDE | 3 |
| 2019 | CATS: Cross-Platform E-Commerce Fraud DetectionabstractNowadays, the popularity of e-commerce has brought huge economic benefits to factories, third-party merchants, and e-commerce service providers. Driven by such huge economic benefits, malicious merchants attempt to promote items through inserting fraudulent purchases, fake review scores, and/or feedback, into them. Mitigating this threat is challenging due to the difficulty of obtaining internal e-commerce data, the variance of e-commerce services used by malicious merchants, and the reluctance of service providers in cooperation. In this paper, we present an efficient, platform-independent, and robust e-commerce fraud detection system, CATS, to detect frauds for different large-scale e-commerce platforms. We implement the design of CATS into a prototype system and evaluate this prototype on the world's popular e-commerce platform Taobao. The evaluation result on Taobao shows that CATS can achieve a high accuracy of 91% in detecting frauds. Based on this success, we then apply CATS on another large-scale e-commerce platforms, and again CATS achieves an accuracy of 96%, suggesting that CATS is very effective on real e-commerce platforms. Based on the cross-platform evaluation results, we conduct a comprehensive analysis on the reported frauds and reveal several abnormal yet interesting behaviors of those reported frauds. Our study in this paper is expected to shed light on defending against frauds for various e-commerce platforms. Haiqin Weng, Shouling Ji, Fuzheng Duan, Zhao Li 0007, Jianhai Chen, Qinming He, Ting Wang 0006 |
ICDE | 1 |
| 2018 | Online E-Commerce Fraud: A Large-Scale Detection and AnalysisabstractNowadays, e-commerce has become prevalent world-wide. With the big success of e-commerce, many malicious promotion services also rise: with the goal of increasing sales, malicious merchants attempt to promote their target items by illegally optimizing the search results using fake visits, purchases, etc. In this paper, we study the fraud detection problem on large-scale e-commerce platforms. First, we develop an efficient and scalable AnTi-Fraud system (ATF) to detect e-commerce frauds for large-scale e-commerce platforms, and implement it in parallel on a large-scale computing platform, called Open Data Processing Service (ODPS). Then, we evaluate ATF using two real large-scale e-commerce datasets (with tens of millions users and items). The results demonstrate that both the precision and the recall of ATF can achieve 0.97+, which suggests that ATF is very effective. More importantly, we deploy ATF on the Taobao platform of Alibaba, which is one of the world's largest e-commerce platforms. The evaluation results show that ATF can also achieve an accuracy of 98.16% on Taobao, which again suggests that ATF is very effective and deployable in practice. Our study in this paper is expected to shed light on defending against online frauds for practical e-commerce platforms. Haiqin Weng, Zhao Li 0007, Shouling Ji, Chen Chu, Haifeng Lu, Tianyu Du, Qinming He |
ICDE | 1 |
| 2018 | Rare category exploration with noisy labels
Haiqin Weng, Kevin Chiew, Zhenguang Liu, Qinming He, Roger Zimmermann |
Expert Syst. Appl. | 1 |
| 2015 | Rare Category Detection ForestabstractRare category detecion (RCD) aims to discover rare categories in a massive unlabeled data set with the help of a labeling oracle. A challenging task in RCD is to discover rare categories which are concealed by numerous data examples from major categories. Only a few algorithms have been proposed for this issue, most of which are on quadratic or cubic time complexity. In this paper, we propose a novel tree-based algorithm known as RCD-Forest with $$O(\varphi n \log {(n/s)})$$ time complexity and high query efficiency where n is the size of the unlabeled data set. Experimental results on both synthetic and real data sets verify the effectiveness and efficiency of our method. Haiqin Weng, Zhenguang Liu, Kevin Chiew, Qinming He |
KSEM | 1 |