EDBT 2026 Demo / reviewers in the wild / expert
Baolin Zheng
dblp:276/3969
· DBLP profile ↗
15ranked-venue papers
2as first author
15since 2021 · last 2026
0009-0002-0381-0255ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 1 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Security and privacy · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language ModelsabstractBaolin Zheng, Guanlin Chen, Qingyang Teng, Hongqiong Zhong, Yingshui Tan, Zhendong Liu, Weixun Wang, Jiaheng Liu, Jian Yang, Huiyun Jing, Jincheng Wei, Wenbo Su, Xiaoyong Zhu, Bo Zheng, Kaifu Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Baolin Zheng, Qingyang Teng, Hongqiong Zhong, Yingshui Tan, Weixun Wang, Jian Yang 0037, Huiyun Jing, Jincheng Wei, Wenbo Su, Xiaoyong Zhu, Bo Zheng 0007, Kaifu Zhang |
ACL (1) | 1 |
| 2025 | Towards Visualization-of-Thought Jailbreak Attack against Large Visual Language ModelsabstractAs Visual Language Models (VLMs) continue to evolve, they have demonstrated increasingly sophisticated logical reasoning capabilities and multimodal thought generation, opening doors to widespread applications. However, this advancement raises serious concerns about content security, particularly when these models process complex multimodal inputs requiring intricate reasoning. When faced with these safety challenges, the critical competition between logical reasoning and safety objectives of VLMs is often overlooked in previous works. In this paper, we introduce Visualization-of-Thought Attack (\textbf{VoTA}), a novel and automated attack framework that strategically constructs chains of images with risky visual thoughts to challenge victim models. Our attack provokes the inherent conflict between the model's logical processing and safety protocols, ultimately leading to the generation of unsafe content. Through comprehensive experiments, VoTA achieves remarkable effectiveness, improving the average attack success rate (ASR) by 26.71\% (from 63.70\% to 90.41\%) on 9 open-source and 6 commercial VLMs, compared to the state-of-the-art methods. These results expose a critical vulnerability: current VLMs struggle to maintain safety guarantees when processing insecure multimodal visualization-of-thought inputs, highlighting the urgency and necessity of enhancing safety alignment. Our code and dataset are available at
https://github.com/Hongqiong12/VoTA.
Content Warning: This paper contains harmful contents that may be offensive. Hongqiong Zhong, Qingyang Teng, Baolin Zheng, Yingshui Tan, Wenbo Su, Xiaoyong Zhu, Bo Zheng 0007, Kaifu Zhang |
NeurIPS | 3 |
| 2025 | Exploring black-box adversarial attacks on Interpretable Deep Learning Systems
Yike Zhan, Baolin Zheng, Dongxin Liu, Boren Deng |
Comput. Vis. Image Underst. | 2 |
| 2024 | From Toxic to Trustworthy: Using Self-Distillation and Semi-supervised Methods to Refine Neural NetworksabstractDespite the tremendous success of deep neural networks (DNNs) across various fields, their susceptibility to potential backdoor attacks seriously threatens their application security, particularly in safety-critical or security-sensitive ones. Given this growing threat, there is a pressing need for research into purging backdoors from DNNs. However, prior efforts on erasing backdoor triggers not only failed to withstand increasingly powerful attacks but also resulted in reduced model performance. In this paper, we propose From Toxic to Trustworthy (FTT), an innovative approach to eliminate backdoor triggers while simultaneously enhancing model accuracy. Following the stringent and practical assumption of limited availability of clean data, we introduce a self-attention distillation (SAD) method to remove the backdoor by aligning the shallow and deep parts of the network. Furthermore, we first devise a semi-supervised learning (SSL) method that leverages ubiquitous and available poisoned data to further purify backdoors and improve accuracy. Extensive experiments on various attacks and models have shown that our FTT can reduce the attack success rate from 97% to 1% and improve the accuracy of 4% on average, demonstrating its effectiveness in mitigating backdoor attacks and improving model performance. Compared to state-of-the-art (SOTA) methods, our FTT can reduce the attack success rate by 2 times and improve the accuracy by 5%, shedding light on backdoor cleansing. Baolin Zheng, Jianbao Hu, Chengyang Li 0001, Xiaoying Bai |
AAAI | 2 |
| 2024 | FIBA: Federated Invisible Backdoor AttackabstractAlthough previous studies have proposed backdoor attacks in Federated Learning (FL), the introduced triggers in these works are easily detectable by human eyes. Besides, this paper also shows that traditional invisible centralized backdoor attacks struggle to work well in FL scenarios. To address this issue, we propose the Federated Invisible Backdoor Attack (FIBA), a novel approach to invisible backdoor attacks against FL. FIBA can balance effectiveness and stealthiness through the auto-adjusted Quality-of-Experience (QoE) restriction and attention-based L2 regularization. It further enhances durability by targeting stable neural connections. Experimental results indicate FIBA’s superior effectiveness, comparable stealthiness, and increased durability over existing methods, demonstrating its ability to bypass current detection mechanisms and robust aggregation techniques in FL. Our codes are available here1. Baolin Zheng |
ICASSP | 2 |
| 2024 | FastTextDodger: Decision-Based Adversarial Attack Against Black-Box NLP Models With Extremely High EfficiencyabstractRecently, achieving query-efficient adversarial example attacks targeting black-box natural language models has attracted widespread attention from researchers. This task is considered difficult due to the discrete nature of texts, limited knowledge of the target model, and strict query access limitations in real-world systems. However, existing attacks often require a large number of queries or result in low attack success rates, having not met practical requirements. To address this, we propose FastTextDodger, a simple and compact decision-based black-box textual adversarial attack that generates grammatically correct adversarial texts with high attack success rates and few queries. Experimental results show that FastTextDodger achieves an impressive 97.4% attack success rate on benchmark datasets and models, and only needs about 200 queries. Compared to state-of-the-art attacks, FastTextDodger only requires one-tenth of the number of queries in text classification and entailment tasks while maintaining comparable attack success rates and perturbed word rates. Xiaoxue Hu, Geling Liu, Baolin Zheng, Lingchen Zhao, Qian Wang 0002, Minxin Du |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2024 | Perception-Driven Imperceptible Adversarial Attack Against Decision-Based Black-Box ModelsabstractAdversarial examples (AEs) pose significant threats to deep neural networks (DNNs), as they can deceive models into making incorrect predictions through craftily-designed malicious perturbations. The emergence of decision-based attacks, which rely solely on the top-1 decision label, further increases risks for real-world black-box models. Currently, the prevailing practice for generating effective AEs in decision-based attacks involves penalizing adversarial perturbations using the ℓp-norm. However, this approach often fails to consider the human perception of adversarial perturbations in real-world scenarios. To tackle this issue, we propose a novel and efficient Imperceptible Decision-based Black-box Attack (IDBA). Our method prioritizes optimizing the perception-related distribution of perturbations, rather than solely focusing on the ℓp-norm. Specifically, IDBA analyzes the perceptual preferences of both models and the human vision system, selectively perturbing components that influence model decisions yet remain imperceptible to human eyes. Extensive experiments demonstrate the superior performance of IDBA in both invisibility and query efficiency, a widely used metric in prior works, in comparison to state-of-the-art methods. With only 4.8K queries, IDBA achieves a Feature SIMilarity (FSIM) score of 0.92 while reducing the Learned Perceptual Image Patch Similarity (LPIPS) to 0.12, indicating remarkable imperceptibility. Shenyi Zhang, Baolin Zheng, Peipei Jiang 0002, Lingchen Zhao, Chao Shen 0001, Qian Wang 0002 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2023 | Adversarial Network Pruning by Filter Robustness EstimationabstractNetwork pruning has been extensively studied in model compression to reduce neural networks’ memory, latency, and computation cost. However, the pruned networks still suffer from the threat posed by adversarial examples, limiting the broader application of the pruned networks in safety-critical applications. Previous studies maintain the robustness of the pruned networks by combining adversarial training and network pruning but ignore preserving the robustness at a high sparsity ratio in structured pruning. To address such a problem, we propose an effective filter importance criterion, Filter Robustness Estimation (FRE), to evaluate the importance of filters by estimating their contribution to the adversarial training loss. Empirical results show that our FRE-based Robustness-aware Filter Pruning (FRFP) outperforms the state-of-the-art methods by 12.19%∼37.01% of empirical robust accuracy on the CIFAR10 dataset with the VGG16 network at an extreme pruning ratio of 90%. Xinlu Zhuang, Yunjie Ge, Baolin Zheng, Qian Wang 0002 |
ICASSP | 3 |
| 2023 | SDBC: A Novel and Effective Self-Distillation Backdoor Cleansing ApproachabstractDeep Neural Networks (DNNs) are vulnerable to backdoor attacks, which only need to poison a small portion of samples to control the behavior of the target model. Moreover, the escalating stealth and power of backdoor attacks present not only significant challenges to backdoor defenses but also enormous potential threats to the widespread adoption of DNNs. In this paper, we propose a novel backdoor defense framework, called Self-Distillation Backdoor Cleansing (SDBC), to remove backdoor triggers from the attacked model. For the practical scenario where only a very small portion of clean data is available, SDBC first introduces self-distillation to clean the backdoor in DNNs. Extensive experiments demonstrate that SDBC can effectively remove backdoor triggers under 6 state-of-the-art backdoor attacks using less than 5% or even less than 1% clean training data without compromising accuracy. Experimental results show that the proposed SDBC outperforms existing state-of-the-art (SOTA) methods, reducing the average ASR from 95.36% to 5.75% and increasing the average ACC by 1.92%. Sheng Ran, Baolin Zheng |
ICONIP (12) | 2 |
| 2023 | Sequence As Genes: An User Behavior Modeling Framework for Fraud Transaction Detection in E-commerceabstractWith the explosive growth of e-commerce, detecting fraudulent transactions in real-world scenarios is becoming increasingly important for e-commerce platforms. Recently, several supervised approaches have been proposed to use user behavior sequences, which record the user's track on platforms and contain rich information for fraud transaction detection. Nevertheless, these methods always suffer from the scarcity of labeled data in real-world scenarios. The recent remarkable pre-training methods in Natural Language Processing (NLP) and Computer Vision (CV) domains offered glimmers of light. However, user behavior sequences differ intrinsically from text, images, and videos. In this paper, we propose a novel and general user behavior pre-training framework, named Sequence As GEnes (SAGE), which provides a new perspective for user behavior modeling. Following the inspiration of treating sequences as genes, we carefully designed the user behavior data organization paradigm and pre-training scheme. Specifically, we propose an efficient data organization paradigm inspired by the nature of DNA expression, which decouples the length of behavior sequences and the corresponding time spans. Also inspired by the natural mechanisms in genetics, we propose two pre-training tasks, namely sequential mutation and sequential recombination, to improve the robustness and consistency of user behavior representations in complicated real-world scenes. Extensive experiments on four differentiated fraud transaction detection real scenarios demonstrate the effectiveness of our proposed framework. Qianru Wu, Baolin Zheng, Yanjie Shi |
KDD | 3 |
| 2022 | Towards Black-Box Adversarial Attacks on Interpretable Deep Learning SystemsabstractRecent works have empirically shown that neural network interpretability is susceptible to malicious manipulations. However, existing attacks against Interpretable Deep Learning Systems (IDLSes) all focus on the white-box setting, which is obviously unpractical in real-world scenarios. In this paper, we make the first attempt to attack IDLSes in the decision-based black-box setting. We propose a new framework called Dual Black-box Adversarial Attack (DBAA) which can generate adversarial examples that are misclassified as the target class, yet have very similar interpretations to their benign cases. We conduct comprehensive experiments on different combinations of classifiers and interpreters to illustrate the effectiveness of DBAA. Empirical results show that in all the cases, DBAA achieves high attack success rates and Intersection over Union (IoU) scores. Yike Zhan, Baolin Zheng, Qian Wang 0002, Ningping Mou, Binqing Guo, Qi Li 0002, Chao Shen 0001, Cong Wang 0001 |
ICME | 2 |
| 2022 | A Few Seconds Can Change Everything: Fast Decision-based Attacks against DNNsabstractPrevious researches have demonstrated deep learning models' vulnerabilities to decision-based adversarial attacks, which craft adversarial examples based solely on information from output decisions (top-1 labels). However, existing decision-based attacks have two major limitations, i.e., expensive query cost and being easy to detect. To bridge the gap and enlarge real threats to commercial applications, we propose a novel and efficient decision-based attack against black-box models, dubbed FastDrop, which only requires a few queries and work well under strong defenses. The crux of the innovation is that, unlike existing adversarial attacks that rely on gradient estimation and additive noise, FastDrop generates adversarial examples by dropping information in the frequency domain. Extensive experiments on three datasets demonstrate that FastDrop can escape the detection of the state-of-the-art (SOTA) black-box defenses and reduce the number of queries by 13~133× under the same level of perturbations compared with the SOTA attacks. FastDrop only needs 10~20 queries to conduct an attack against various black-box models within 1s. Besides, on commercial vision APIs provided by Baidu and Tencent, FastDrop achieves an attack success rate (ASR) of 100% with 10 queries on average, which poses a real and severe threat to real-world applications. Ningping Mou, Baolin Zheng, Qian Wang 0002, Yunjie Ge, Binqing Guo |
IJCAI | 2 |
| 2021 | Black-box Adversarial Attacks on Commercial Speech Platforms with Minimal InformationabstractAdversarial attacks against commercial black-box speech platforms, including cloud speech APIs and voice control devices, have received little attention until recent years. Constructing such attacks is difficult mainly due to the unique characteristics of time-domain speech signals and the much more complex architecture of acoustic systems. The current "black-box" attacks all heavily rely on the knowledge of prediction/confidence scores or other probability information to craft effective adversarial examples (AEs), which can be intuitively defended by service providers without returning these messages. In this paper, we take one more step forward and propose two novel adversarial attacks in more practical and rigorous scenarios. For commercial cloud speech APIs, we propose Occam, a decision-only black-box adversarial attack, where only final decisions are available to the adversary. In Occam, we formulate the decision-only AE generation as a discontinuous large-scale global optimization problem, and solve it by adaptively decomposing this complicated problem into a set of sub-problems and cooperatively optimizing each one. Our Occam is a one-size-fits-all approach, which achieves 100% success rates of attacks (SRoA) with an average SNR of 14.23dB, on a wide range of popular speech and speaker recognition APIs, including Google, Alibaba, Microsoft, Tencent, iFlytek, and Jingdong, outperforming the state-of-the-art black-box attacks. For commercial voice control devices, we propose NI-Occam, the first non-interactive physical adversarial attack, where the adversary does not need to query the oracle and has no access to its internal information and training data. We, for the first time, combine adversarial attacks with model inversion attacks, and thus generate the physically-effective audio AEs with high transferability without any interaction with target devices. Our experimental results show that NI-Occam can successfully fool Apple Siri, Microsoft Cortana, Google Assistant, iFlytek and Amazon Echo with an average SRoA of 52% and SNR of 9.65dB, shedding light on non-interactive physical attacks against voice control devices. Baolin Zheng, Peipei Jiang 0002, Qian Wang 0002, Qi Li 0002, Chao Shen 0001, Cong Wang 0001, Yunjie Ge, Qingyang Teng, Shenyi Zhang |
CCS | 1 |
| 2021 | Anti-Distillation Backdoor Attacks: Backdoors Can Really Survive in Knowledge DistillationabstractMotivated by resource-limited scenarios, knowledge distillation (KD) has received growing attention, effectively and quickly producing lightweight yet high-performance student models by transferring the dark knowledge from large teacher models. However, many pre-trained teacher models are downloaded from public platforms that lack necessary vetting, posing a possible threat to knowledge distillation tasks. Unfortunately, thus far, there has been little research to consider the backdoor attack from the teacher model into student models in KD, which may pose a severe threat to its wide use. In this paper, we, for the first time, propose a novel Anti-Distillation Backdoor Attack (ADBA), in which the backdoor embedded in the public teacher model can survive the knowledge distillation process and thus be transferred to secret distilled student models. We first introduce a shadow to imitate the distillation process and adopt an optimizable trigger to transfer information to help craft the desired teacher model. Our attack is powerful and effective, which achieves 95.92%, 94.79%, and 90.19% average success rates of attacks (SRoAs) against several different structure student models on MNIST, CIFAR-10, and GTSRB, respectively. Our ADBA also performs robustly under different user distillation environments with 91.72% and 92.37% average SRoAs on MNIST and CIFAR-10, respectively. Finally, we show that the ADBA has a low overhead in the injecting process, which converges on 50 and 70 epochs on CIFAR-10 and GTSRB, respectively, while the normal training epochs of these datasets are almost 200. Yunjie Ge, Qian Wang 0002, Baolin Zheng, Xinlu Zhuang, Qi Li 0002, Chao Shen 0001, Cong Wang 0001 |
ACM Multimedia | 3 |
| 2021 | Towards Query-Efficient Adversarial Attacks Against Automatic Speech Recognition SystemsabstractAdversarial attacks, which attract explosive rese- arch attention in recent years, have achieved fantastic success in fooling neural networks, especially for image-classification tasks. While for automatic speech recognition (ASR) tasks, the state-of-the-arts mainly focus on white-box attacks where the adversary is assumed to get full access to the details inside the system, e.g., network architecture, weights, etc. However, this assumption does not hold in practice. The construction of real-world adversarial examples against ASR systems is still a very challenging problem. In this paper, we, for the first time, present a novel and effective attack on ASR systems, named Selective Gradient Estimation Attack (SGEA). Compared with prior literatures, SGEA only needs limited access to the output probabilities of neural networks, and achieves extremely high efficiency and success rates. We attacked the DeepSpeech system on Mozilla Common Voice and LibriSpeech datasets in our experiments. The results demonstrate that SGEA improves the attack success rate from 35% to 98%, while reducing the number of queries by 66%. Qian Wang 0002, Baolin Zheng, Qi Li 0002, Chao Shen 0001, Zhongjie Ba |
IEEE Trans. Inf. Forensics Secur. | 2 |