EDBT 2026 Demo / reviewers in the wild / expert
Zhicong Yan
dblp:262/5822
· DBLP profile ↗
5ranked-venue papers
3as first author
5since 2021 · last 2023
0000-0003-3363-7713ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 3 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | DHBE: Data-free Holistic Backdoor Erasing in Deep Neural Networks via Restricted Adversarial DistillationabstractBackdoor attacks have emerged as an urgent threat to Deep Neural Networks (DNNs), where victim DNNs are furtively implanted with malicious neurons that could be triggered by the adversary. To defend against backdoor attacks, many works establish a staged pipeline to remove backdoors from victim DNNs: inspecting, locating, and erasing. However, in a scenario where a few clean data can be accessible, such pipeline is fragile and cannot erase backdoors completely without sacrificing model accuracy. To address this issue, in this paper, we propose a novel data-free holistic backdoor erasing (DHBE) framework. Instead of the staged pipeline, the DHBE treats the backdoor erasing task as a unified adversarial procedure, which seeks equilibrium between two different competing processes: distillation and backdoor regularization. In distillation, the backdoored DNN is distilled into a proxy model, transferring its knowledge about clean data, yet backdoors are simultaneously transferred. In backdoor regularization, the proxy model is holistically regularized to prevent from infecting any possible backdoor transferred from distillation. These two processes jointly proceed with data-free adversarial optimization until a clean, high-accuracy proxy model is obtained. With the novel adversarial design, our framework demonstrates its superiority in three aspects: 1) minimal detriment to model accuracy, 2) high tolerance for hyperparameters, and 3) no demand for clean data. Extensive experiments on various backdoor attacks and datasets are performed to verify the effectiveness of the proposed framework. Code is available at https://github.com/yanzhicong/DHBE Zhicong Yan, Shenghong Li 0001, Ruijie Zhao 0001, Yuan Tian 0017 |
AsiaCCS | 1 |
| 2023 | GeeSolver: A Generic, Efficient, and Effortless Solver with Self-Supervised Learning for Breaking Text CaptchasabstractAlthough text-based captcha, which is used to differentiate between human users and bots, has faced many attack methods, it remains a widely used security mechanism and is employed by some websites. Some deep learning-based text captcha solvers have shown excellent results, but the labor-intensive and time-consuming labeling process severely limits their viability. Previous works attempted to create easy-to-use solvers using a limited collection of labeled data. However, they are hampered by inefficient preprocessing procedures and inability to recognize the captchas with complicated security features.In this paper, we propose GeeSolver, a generic, efficient, and effortless solver for breaking text-based captchas based on self-supervised learning. Our insight is that numerous difficult-to-attack captcha schemes that "damage" the standard font of characters are similar to image masks. And we could leverage masked autoencoders (MAE) to improve the captcha solver to learn the latent representation from the "unmasked" part of the captcha images. Specifically, our model consists of a ViT encoder as latent representation extractor and a well-designed decoder for captcha recognition. We apply MAE paradigm to train our encoder, which enables the encoder to extract latent representation from local information (i.e., without masking part) that can infer the corresponding character. Further, we freeze the parameters of the encoder and leverage a few labeled captchas and many unlabeled captchas to train our captcha decoder with semi-supervised learning.Our experiments with real-world captcha schemes demonstrate that GeeSolver outperforms the state-of-the-art methods by a large margin using a few labeled captchas. We also show that GeeSolver is highly efficient as it can solve a captcha within 25 ms using a desktop CPU and 9 ms using a desktop GPU. Besides, thanks to latent representation extraction, we successfully break the hard-to-attack captcha schemes, proving the generality of our solver. We hope that our work will help security experts to revisit the design and availability of text-based captchas. The code is available at https://github.com/NSSL-SJTU/GeeSolver. Ruijie Zhao 0001, Xianwen Deng, Zhicong Yan, Zhengguang Han, Libo Chen 0001, Zhi Xue |
SP | 4 |
| 2022 | MT-FlowFormer: A Semi-Supervised Flow Transformer for Encrypted Traffic ClassificationabstractWith the increasing demand for the protection of personal network meta-data, encrypted networks have grown in popularity, so do the challenge of monitoring and analyzing encrypted network traffic. Currently, some deep learning-based methods have been proposed to leverage statistical features for encrypted traffic classification, which are barely affected by encryption techniques. However, these works still suffer from two main intrinsic limitations: (1) the feature extraction process lacks a mechanism to take into account correlations between flows in the flow sequence; and (2) a large volume of manually-labeled data is required for training an effective deep classifier. In this paper, we propose a novel semi-supervised framework to address these problems. To be specific, an efficient classifier with attention mechanism is proposed to extract features from flow sequences with low computational cost. Then, a Mean Teacher-style semi-supervised framework is adopted to exploit the unlabeled traffic data, where a spatiotemporal data augmentation method is designed as the key component to explore the spatial and temporal relationship within the unlabeled traffic data. Experimental results on two real-world traffic datasets demonstrate that our method outperforms state-of-the-art methods with a large margin. Ruijie Zhao 0001, Xianwen Deng, Zhicong Yan, Zhi Xue |
KDD | 3 |
| 2021 | DeHiB: Deep Hidden Backdoor Attack on Semi-supervised Learning via Adversarial PerturbationabstractThe threat of data-poisoning backdoor attacks on learning algorithms typically comes from the labeled data. However, in deep semi-supervised learning (SSL), unknown threats mainly stem from the unlabeled data. In this paper, we propose a novel deep hidden backdoor (DeHiB) attack scheme for SSL-based systems. In contrast to the conventional attacking methods, the DeHiB can inject malicious unlabeled training data to the semi-supervised learner so as to enable the SSL model to output premeditated results. In particular, a robust adversarial perturbation generator regularized by a unified objective function is proposed to generate poisoned data. To alleviate the negative impact of the trigger patterns on model accuracy and improve the attack success rate, a novel contrastive data poisoning strategy is designed. Using the proposed data poisoning scheme, one can implant the backdoor into the SSL model using the raw data without hand-crafted labels. Extensive experiments based on CIFAR10 and CIFAR100 datasets demonstrated the effectiveness and crypticity of the proposed scheme. Zhicong Yan, Gaolei Li, Yuan Tian 0017, Jun Wu 0001, Shenghong Li 0001, Mingzhe Chen, H. Vincent Poor |
AAAI | 1 |
| 2021 | Deep Neural Backdoor in Semi-Supervised Learning: Threats and CountermeasuresabstractSemi-Supervised Learning (SSL) is a powerful derivative for humans to discover the hidden knowledge, and will be a great substitute for data taggers. Although the availability of unlabeled data rises up a huge passion to SSL, the untrustness of unlabeled data leads to many unknown security risks. In this paper, we first identify an insidious backdoor threat of SSL where unlabeled training data are poisoned by backdoor methods migrated from supervised settings. Then, to further exploit this threat, a Deep Neural Backdoor (DeNeB) scheme is proposed, which requires less data poisoning budgets and produces stronger backdoor effectiveness. By poisoning a fraction of unlabeled training data, the DeNeB achieves the illegal manipulation on the trained model without modifying the training process. Finally, an efficient detection-and-purification defense (DePuD) framework is proposed to thwart the proposed scheme. In DePuD, we construct a deep detector to locate trigger patterns in the unlabeled training data, and perform secured SSL training with purified unlabeled data where the detected trigger patterns are obfuscated. Extensive experiments based on benchmark datasets are performed to demonstrate the huge threatening of DeNeB and the effectiveness of DePuD. To our best knowledge, this is the first work to achieve the backdoor and its defense in semi-supervised learning. Zhicong Yan, Jun Wu 0001, Gaolei Li, Shenghong Li 0001, Mohsen Guizani |
IEEE Trans. Inf. Forensics Secur. | 1 |