VLDB 2026 Research / reviewers in the wild / expert
Bingquan Shen
dblp:151/9308
· DBLP profile ↗
9ranked-venue papers
1as first author
8since 2021 · last 2025
0009-0006-6442-551XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Security and privacy · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Towards Effective and Robust Unlearnable Examples Against Object DetectionabstractObject detection has become crucial due to its extensive applications across various industries. However, the data used to train these models is often sensitive and proprietary, raising significant concerns about its security and unauthorized usage. Unlearnable examples (UEs) represent a promising strategy to safeguard proprietary datasets by embedding imperceptible perturbations that degrade model performance when such data is used during training. This paper explores UEs tailored specifically for object detection tasks, which pose unique challenges due to the multi-task nature of object detection. We propose a novel framework that generates robust and effective UEs, and significantly degrades object detector performance while maintaining imperceptibility. Comprehensive experiments demonstrate the resilience of the proposed UEs against various countermeasures, underscoring their potential as a practical solution for protecting data in object detection. Chenyu Yi, Ruohan Meng, Haohang Peng, Bingquan Shen, Alex Chichung Kot |
ICIP | 4 |
| 2025 | Defending Multimodal Backdoored Models by Repulsive Visual Prompt TuningabstractMultimodal contrastive learning models (e.g., CLIP) can learn high-quality representations from large-scale image-text datasets, while they exhibit significant vulnerabilities to backdoor attacks, raising serious safety concerns. In this paper, we reveal that CLIP's vulnerabilities primarily stem from its tendency to encode features beyond in-dataset predictive patterns, compromising its visual feature resistivity to input perturbations. This makes its encoded features highly susceptible to being reshaped by backdoor triggers. To address this challenge, we propose Repulsive Visual Prompt Tuning (RVPT), a novel defense approach that employs deep visual prompt tuning with a specially designed feature-repelling loss. Specifically, RVPT adversarially repels the encoded features from deeper layers while optimizing the standard cross-entropy loss, ensuring that only predictive features in downstream tasks are encoded, thereby enhancing CLIP’s visual feature resistivity against input perturbations and mitigating its susceptibility to backdoor attacks. Unlike existing multimodal backdoor defense methods that typically require the availability of poisoned data or involve fine-tuning the entire model, RVPT leverages few-shot downstream clean samples and only tunes a small number of parameters. Empirical results demonstrate that RVPT tunes only 0.27\% of the parameters in CLIP, yet it significantly outperforms state-of-the-art defense methods, reducing the attack success rate from 89.70\% to 2.76\% against the most advanced multimodal attacks on ImageNet and effectively generalizes its defensive capabilities across multiple datasets. Our code is available on https://anonymous.4open.science/r/rvpt-anonymous. Zhifang Zhang, Shuo He 0001, Haobo Wang 0001, Bingquan Shen, Lei Feng 0006 |
NeurIPS | 4 |
| 2024 | Flexible-Modal Deception Detection with Audio-Visual AdapterabstractDeception detection within audio-visual modalities is vital across diverse sectors, notably in customs security and multimedia anti-fraud. However, this notable efficacy is lost by the necessity to train and deploy separate models for each conceivable modality scenario, leading to redundancy and inefficiency. Moreover, real-world environments where multi-modal models are deployed often fail to meet these idealized conditions. To overcome these challenges and further elevate performance levels, we propose an advanced Transformer-based framework complemented by an Audio-Visual Adapter (AVA) integrating temporal features from both audio and visual modalities. In addition, we introduce an innovative multi-modal contrastive learning method that is designed to enhance the correlation between uni-modal features and their integrated counterparts within a consistent feature space. Our designed method can deal with the flexible-model scenario instead of deploying different models for various modalities. Empirical evaluations conducted on two benchmark datasets have validated the superiority of our proposed model over other multi-modal fusion techniques, particularly in scenarios characterized by varying and missing modalities. This strongly affirms the effectiveness of our approach in significantly boosting the accuracy of deception detection in complex, real-world multi-modal scenarios. The codes will be released soon. Zhaoxu Li, Zitong Yu, Xun Lin, Nithish Muthuchamy Selvaraj, Xiaobao Guo, Bingquan Shen, Adams Wai-Kin Kong, Alex Chichung Kot |
IJCB | 6 |
| 2024 | Backdoor Secrets Unveiled: Identifying Backdoor Data with Optimized Scaled Prediction ConsistencyabstractModern machine learning (ML) systems demand substantial training data, often resorting to external sources. Nevertheless, this practice renders them vulnerable to backdoor poisoning attacks. Prior backdoor defense strategies have primarily focused on the identification of backdoored models or poisoned data characteristics, typically operating under the assumption of access to clean data. In this work, we delve into a relatively underexplored challenge: the automatic identification of backdoor data within a poisoned dataset, all under realistic conditions, *i.e.*, without the need for additional clean data or without manually defining a threshold for backdoor detection. We draw an inspiration from the scaled prediction consistency (SPC) technique, which exploits the prediction invariance of poisoned data to an input scaling factor. Based on this, we pose the backdoor data identification problem as a hierarchical data splitting optimization problem, leveraging a novel SPC-based loss function as the primary optimization objective. Our innovation unfolds in several key aspects. First, we revisit the vanilla SPC method, unveiling its limitations in addressing the proposed backdoor identification problem. Subsequently, we develop a bi-level optimization-based approach to precisely identify backdoor data by minimizing the advanced SPC loss. Finally, we demonstrate the efficacy of our proposal against a spectrum of backdoor attacks, encompassing basic label-corrupted attacks as well as more sophisticated clean-label attacks, evaluated across various benchmark datasets. Experiment results show that our approach often surpasses the performance of current baselines in identifying backdoor data points, resulting in about 4\%-36\% improvement in average AUROC. Codes are available at https://github.com/OPTML-Group/BackdoorMSPC. Soumyadeep Pal, Yuguang Yao, Ren Wang 0008, Bingquan Shen, Sijia Liu 0001 |
ICLR | 4 |
| 2024 | From Trojan Horses to Castle Walls: Unveiling Bilateral Data Poisoning Effects in Diffusion ModelsabstractWhile state-of-the-art diffusion models (DMs) excel in image generation, concerns regarding their security persist. Earlier research highlighted DMs' vulnerability to data poisoning attacks, but these studies placed stricter requirements than conventional methods like 'BadNets' in image classification. This is because the art necessitates modifications to the diffusion training and sampling procedures. Unlike the prior work, we investigate whether BadNets-like data poisoning methods can directly degrade the generation by DMs. In other words, if only the training dataset is contaminated (without manipulating the diffusion process), how will this affect the performance of learned DMs? In this setting, we uncover bilateral data poisoning effects that not only serve an adversarial purpose (compromising the functionality of DMs) but also offer a defensive advantage (which can be leveraged for defense in classification tasks against poisoning attacks). We show that a BadNets-like data poisoning attack remains effective in DMs for producing incorrect images (misaligned with the intended text conditions). Meanwhile, poisoned DMs exhibit an increased ratio of triggers, a phenomenon we refer to as 'trigger amplification', among the generated images. This insight can be then used to enhance the detection of poisoned training data. In addition, even under a low poisoning ratio, studying the poisoning effects of DMs is also valuable for designing robust image classifiers against such attacks. Last but not least, we establish a meaningful linkage between data poisoning and the phenomenon of data replications by exploring DMs' inherent data memorization tendencies. Code is available at https://github.com/OPTML-Group/BiBadDiff. Zhuoshi Pan, Yuguang Yao, Gaowen Liu, Bingquan Shen, H. Vicky Zhao, Ramana Rao Kompella, Sijia Liu 0001 |
NeurIPS | 4 |
| 2024 | Semantic Deep Hiding for Robust Unlearnable ExamplesabstractEnsuring data privacy and protection has become paramount in the era of deep learning. Unlearnable examples are proposed to mislead the deep learning models and prevent data from unauthorized exploration by adding small perturbations to data. However, such perturbations (e.g., noise, texture, color change) predominantly impact low-level features, making them vulnerable to common countermeasures. In contrast, semantic images with intricate shapes have a wealth of high-level features, making them more resilient to countermeasures and potential for producing robust unlearnable examples. In this paper, we propose a Deep Hiding (DH) scheme that adaptively hides semantic images enriched with high-level features. We employ an Invertible Neural Network (INN) to invisibly integrate predefined images, inherently hiding them with deceptive perturbations. To enhance data unlearnability, we introduce a Latent Feature Concentration module, designed to work with the INN, regularizing the intra-class variance of these perturbations. To further boost the robustness of unlearnable examples, we design a Semantic Images Generation module that produces hidden semantic images. By utilizing similar semantic information, this module generates similar semantic images for samples within the same classes, thereby enlarging the inter-class distance and narrowing the intra-class distance. Extensive experiments on CIFAR-10, CIFAR-100, and an ImageNet subset, against 18 countermeasures, reveal that our proposed method exhibits outstanding robustness for unlearnable examples, demonstrating its efficacy in preventing unauthorized data exploitation. Ruohan Meng, Chenyu Yi, Yi Yu 0011, Siyuan Yang 0001, Bingquan Shen, Alex Chichung Kot |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2023 | Audio-Visual Deception Detection: DOLOS Dataset and Parameter-Efficient Crossmodal LearningabstractDeception detection in conversations is a challenging yet important task, having pivotal applications in many fields such as credibility assessment in business, multimedia anti-frauds, and custom security. Despite this, deception detection research is hindered by the lack of high-quality deception datasets, as well as the difficulties of learning multimodal features effectively. To address this issue, we introduce DOLOS1, the largest gameshow deception detection dataset with rich deceptive conversations. DOLOS includes 1, 675 video clips featuring 213 subjects, and it has been labeled with audio-visual feature annotations. We provide train-test, duration, and gender protocols to investigate the impact of different factors. We benchmark our dataset on previously proposed deception detection approaches. To further improve the performance by fine-tuning fewer parameters, we propose Parameter-Efficient Crossmodal Learning (PECL), where a Uniform Temporal Adapter (UT-Adapter) explores temporal attention in transformer-based architectures, and a crossmodal fusion module, Plug-in Audio-Visual Fusion (PAVF), combines crossmodal information from audio-visual features. Based on the rich fine-grained audio-visual annotations on DOLOS, we also exploit multi-task learning to enhance performance by concurrently predicting deception and audiovisual features. Experimental results demonstrate the desired quality of the DOLOS dataset and the effectiveness of the PECL. The DOLOS dataset and the source codes are available at here. Xiaobao Guo, Nithish Muthuchamy Selvaraj, Zitong Yu, Adams Wai-Kin Kong, Bingquan Shen, Alex Chichung Kot |
ICCV | 5 |
| 2023 | Adapter Incremental Continual Learning of Efficient Audio Spectrogram Transformers
Nithish Muthuchamy Selvaraj, Xiaobao Guo, Adams Wai-Kin Kong, Bingquan Shen, Alex Chichung Kot |
INTERSPEECH | 4 |
| 2014 | Functional task based assistance during walking for a Lower Extremity Assistive DeviceabstractIn this paper, we propose a functional task based assistance controller to aid user in the walking task with our Lower Extremity Assistive Device (LEAD). Firstly, a gait period detector, which utilizes a Gaussian Mixture Model (GMM), is developed to estimate the user's current gait period among the six major periods. Then, an impedance based controller is used to apply assistive torques to the hip and knee joints of the user based on the functional task intended at the current gait period. To validate the above control scheme, preliminary experiments have been performed with one healthy subject walking on a treadmill. The results show that the gait period detector can effectively detect each gait period for the whole cycle. Based on measurements of the heart rate, the proposed assistance method has shown that it can effectively assist a human user in walking at speed of 1 km/h. Bingquan Shen, Jinfu Li 0001, Chee-Meng Chew |
ICRA | 1 |