EDBT 2026 Demo / reviewers in the wild / expert
Kai Chen 0027
dblp:181/2839-27
· DBLP profile ↗
11ranked-venue papers
2as first author
9since 2021 · last 2025
0009-0000-5195-6107ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 9 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DuMo: Dual Encoder Modulation Network for Precise Concept ErasureabstractThe exceptional generative capability of text-to-image models has raised substantial safety concerns regarding the generation of Not-Safe-For-Work (NSFW) content and potential copyright infringement. To address these concerns, previous methods safeguard the models by eliminating inappropriate concepts. Nonetheless, these models alter the parameters of the backbone network and exert considerable influences on the structural (low-frequency) components of the image, which undermines the model's ability to retain irrelevant concepts. In this work, we propose our Dual encoder Modulation network (DuMo), which achieves precise erasure of inappropriate target concepts with minimum impairment to non-target concepts. In contrast to previous methods, DuMo employs the Eraser with PRior Knowledge (EPR) module which modifies the skip connection features of the U-NET and primarily achieves concept erasure on details (high-frequency) components of the image. To minimize the demage to non-target concepts during erasure, the parameters of the backbone U-NET are frozen and the prior knowledge from the original skip connection features is introduced to the erasure process. Meanwhile, the phenomenon is observed that distinct erasing preferences for the image structure and details are demonstrated by the EPR at different timesteps and layers. Therefore, we adopt a novel Time-Layer MOdulation process (TLMO) that adjusts the erasure scale of EPR module's outputs across different layers and timesteps, automatically balancing the erasure effects and model's generative ability. Our method achieves state-of-the-art performance on Explicit Content Erasure (detecting only 34 nude parts), Cartoon Concept Removal (with an average LPIPS_da of 0.428, 0.113 higher than SOTA at 0.315), and Artistic Style Erasure (with an average LPIPS_da of 0.387, 0.088 higher than SOTA at 0.299), clearly outperforming alternative methods. Kai Chen 0027, Zhipeng Wei 0001, Jingjing Chen 0001, Yu-Gang Jiang 0001 |
AAAI | 2 |
| 2025 | TAPT: Test-Time Adversarial Prompt Tuning for Robust Inference in Vision-Language ModelsabstractLarge pre-trained Vision-Language Models (VLMs) such as CLIP have demonstrated excellent zero-shot generalizability across various downstream tasks. However, recent studies have shown that the inference performance of CLIP can be greatly degraded by small adversarial perturbations, especially its visual modality, posing significant safety threats. To mitigate this vulnerability, in this paper, we propose a novel defense method called Test-Time Adversarial Prompt Tuning (TAPT) to enhance the inference robustness of CLIP against visual adversarial attacks. TAPT is a test-time defense method that learns defensive bimodal (textual and visual) prompts to robustify the inference process of CLIP. Specifically, it is an unsupervised method that optimizes the defensive prompts for each test sample by minimizing a multi-view entropy and aligning adversarial-clean distributions. We evaluate the effectiveness of TAPT on 11 benchmark datasets, including ImageNet and 10 other zero-shot datasets, demonstrating that it enhances the zero-shot adversarial robustness of the original CLIP by at least 48.9% against AutoAttack (AA), while largely maintaining performance on clean examples. Moreover, TAPT outperforms existing adversarial prompt tuning methods across various backbones, achieving an average robustness improvement of at least 36.6%. Code is available at https://github.com/xinwong/TAPT. Xin Wang 0119, Kai Chen 0027, Jiaming Zhang 0006, Jingjing Chen 0001, Xingjun Ma |
CVPR | 2 |
| 2024 | Reliable and Efficient Concept Erasure of Text-to-Image Diffusion Models
Kai Chen 0027, Zhipeng Wei 0001, Jingjing Chen 0001, Yu-Gang Jiang 0001 |
ECCV (53) | 2 |
| 2024 | AdvQDet: Detecting Query-Based Adversarial Attacks with Adversarial Contrastive Prompt Tuning
Xin Wang 0119, Kai Chen 0027, Xingjun Ma, Zhineng Chen, Jingjing Chen 0001, Yu-Gang Jiang 0001 |
ACM Multimedia | 2 |
| 2024 | ReToMe-VA: Recursive Token Merging for Video Diffusion-based Unrestricted Adversarial AttackabstractRecent diffusion-based unrestricted attacks generate imperceptible adversarial examples with high transferability compared to previous unrestricted attacks and restricted attacks. However, existing works on diffusion-based unrestricted attacks are mostly focused on images yet are seldom explored in videos. In this paper, we propose the Recursive Token Merging for Video Diffusion-based Unrestricted Adversarial Attack (ReToMe-VA), which is the first framework to generate imperceptible adversarial video clips with higher transferability. Specifically, to achieve spatial imperceptibility, ReToMe-VA adopts a Timestep-wise Adversarial Latent Optimization (TALO) strategy that optimizes perturbations in diffusion models' latent space at each denoising step. TALO offers iterative and accurate updates to generate more powerful adversarial frames. TALO can further reduce memory consumption in gradient computation. Moreover, to achieve temporal imperceptibility, ReToMe-VA introduces a Recursive Token Merging (ReToMe) mechanism by matching and merging tokens across video frames in the self-attention module, resulting in temporally consistent adversarial videos. ReToMe concurrently facilitates inter-frame interactions into the attack process, inducing more diverse and robust gradients, thus leading to better adversarial transferability. Extensive experiments demonstrate the efficacy of ReToMe-VA, particularly in surpassing state-of-the-art attacks in adversarial transferability by more than 14.16% on average. Ziyi Gao 0002, Kai Chen 0027, Zhipeng Wei 0001, Tingshu Mou, Jingjing Chen 0001, Zhiyu Tan, Yu-Gang Jiang 0001 |
ACM Multimedia | 2 |
| 2024 | Highly Transferable Diffusion-based Unrestricted Adversarial Attack on Pre-trained Vision-Language ModelsabstractPre-trained Vision-Language Models (VLMs) have shown great ability in various Vision-Language tasks. However, these VLMs exhibit inherent vulnerabilities to transferable adversarial examples, which could potentially undermine their performance and reliability in real-world applications. Cross-modal interactions have been demonstrated to be the key point to boosting adversarial transferability, but the utilization of them is limited in existing multimodal adversarial attacks. Stable Diffusion, which contains multiple cross-attention modules, possesses great potential in facilitating adversarial transferability by leveraging abundant cross-modal interactions. Therefore, We propose a Multimodal Diffusion-based Attack (MDA), which conducts adversarial attacks against VLMs using Stable Diffusion. Specifically, MDA initially generates adversarial text, which is subsequently utilized to optimize the adversarial image during the diffusion process. Besides leveraging adversarial text in calculating downstream loss, MDA also takes it as the guiding prompt in adversarial image generation during the denoising process, which enriches the ways of cross-modal interactions, thus strengthening the adversarial transferability. Compared with pixel-based attacks, MDA introduces perturbations in the latent space rather than pixel space to manipulate high-level semantics, which is also beneficial to improving adversarial transferability. Experimental results demonstrate that the adversarial examples generated by MDA are highly transferable across different VLMs on different downstream tasks, surpassing state-of-the-art methods by a large margin. Wenzhuo Xu, Kai Chen 0027, Ziyi Gao 0002, Zhipeng Wei 0001, Jingjing Chen 0001, Yu-Gang Jiang 0001 |
ACM Multimedia | 2 |
| 2023 | Downstream Task-agnostic Transferable Attacks on Language-Image Pre-training ModelsabstractVision-language pre-trained models (e.g., CLIP) trained on large-scale datasets via self-supervised learning, are drawing increasing research attention since they can achieve superior performances on multi-modal downstream tasks. Nevertheless, we find that the adversarial perturbations crafted on vision-language pre-trained models can be used to attack different corresponding downstream task models. Specifically, to investigate such adversarial transferability, we introduce a task-agnostic method named Global and Local Augmentation (GLA) attack to generate highly transferable adversarial examples on CLIP, to attack black-box downstream task models. GLA adopts random crop and resize at both global and local patch levels, to create more diversity and make adversarial noises robust. Then GLA generates the adversarial perturbations by minimizing the cosine similarity between intermediate features from augmented adversarial and benign examples. Extensive experiments on three CLIP image encoders with different backbones and three different downstream tasks demonstrate the superiority of our method compared with other strong baselines. The code is available at https://github.com/yqlvcoding/GLAattack. Yiqiang Lv, Jingjing Chen 0001, Zhipeng Wei 0001, Kai Chen 0027, Zuxuan Wu, Yu-Gang Jiang 0001 |
ICME | 4 |
| 2023 | GCMA: Generative Cross-Modal Transferable Adversarial Attacks from Images to VideosabstractExisting cross-domain transferable attacks mostly focus on exploring the adversarial transferability across homomodal domains, while the adversarial transferability across heteromodal domains, e.g., image domains to video domains, has received less attention. This paper investigates cross-modal transferable attacks from image domains to video domains with the generator-oriented approach, i.e., crafting adversarial perturbations for each frame of video clips with the perturbation generator trained in the ImageNet domain to attack target video models. To this end, we propose an effective Generative Cross-Modal Attacks (GCMA) framework to enhance adversarial transferability from image domains to video domains. To narrow the domain gap between image and video data, we first propose a random motion module that warps images with synthetic random optical flows. We then integrate the random motion module into the feature disruption loss to incorporate additional temporal cues in the training phase. Specifically, feature disruption loss minimizes the cosine similarity between intermediate features of warped benign and adversarial images. Furthermore, motivated by the positive correlation between transferability and temporal consistency of adversarial video clips, we also introduce a temporal consistency loss that maximizes the cosine similarity between intermediate features of warped adversarial images and adversarial counterparts of warped benign images. Finally, GCMA trains the perturbation generator by simultaneously optimizing feature disruption loss and temporal consistency loss. Extensive experiments demonstrate the effectiveness of our proposed method, achieving state-of-the-art performance on Kinetics-400 and UCF-101. Our code is available at https://github.com/kay-ck/GCMA. Kai Chen 0027, Zhipeng Wei 0001, Jingjing Chen 0001, Zuxuan Wu, Yu-Gang Jiang 0001 |
ACM Multimedia | 1 |
| 2022 | Attacking Video Recognition Models with Bullet-Screen CommentsabstractRecent research has demonstrated that Deep Neural Networks (DNNs) are vulnerable to adversarial patches which introduce perceptible but localized changes to the input. Nevertheless, existing approaches have focused on generating adversarial patches on images, their counterparts in videos have been less explored. Compared with images, attacking videos is much more challenging as it needs to consider not only spatial cues but also temporal cues. To close this gap, we introduce a novel adversarial attack in this paper, the bullet-screen comment (BSC) attack, which attacks video recognition models with BSCs. Specifically, adversarial BSCs are generated with a Reinforcement Learning (RL) framework, where the environment is set as the target model and the agent plays the role of selecting the position and transparency of each BSC. By continuously querying the target models and receiving feedback, the agent gradually adjusts its selection strategies in order to achieve a high fooling rate with non-overlapping BSCs. As BSCs can be regarded as a kind of meaningful patch, adding it to a clean video will not affect people’s understanding of the video content, nor will arouse people’s suspicion. We conduct extensive experiments to verify the effectiveness of the proposed method. On both UCF-101 and HMDB-51 datasets, our BSC attack method can achieve about 90% fooling rate when attacking three mainstream video recognition models, while only occluding < 8% areas in the video. Our code is available at https://github.com/kay-ck/BSC-attack. Kai Chen 0027, Zhipeng Wei 0001, Jingjing Chen 0001, Zuxuan Wu, Yu-Gang Jiang 0001 |
AAAI | 1 |
| 2020 | HMOE-Net: Hybrid Multi-scale Object Equalization Network for Intracerebral Hemorrhage Segmentation in CT ImagesabstractIn this paper, we propose a novel Hybrid Multi-scale Object Equalization Network (HMOE-Net) to segment intracerebral hemorrhage (ICH) regions. In particular, we design a shallow feature extraction network (SFENet) and a deep feature extraction network (DFENet) to solve the problem of equalization learning of hybrid multi-scale object features. The multi-level feature extraction (MLFE) blocks are presented in DFENet to explore multi-level semantic features more effectively. Furthermore, we adopt a progressive feature extraction strategy combining SFENet and DFENet to further consider the differences of various ICH regions and achieve the equalization feature learning of multi-scale objects. To verify the effectiveness of HMOE-Net, we collect a clinical ICH dataset with a total of 500 CT cases from three hospitals for the evaluation. The experimental results show that HMOE-Net is superior to six state-of-the-art methods and achieves accurate segmentation for multi-scale ICH regions. Xizhi He, Kai Chen 0027, Kai Hu 0002, Zhineng Chen, Xuanya Li, Xieping Gao 0001 |
BIBM | 2 |
| 2020 | Automatic segmentation of intracerebral hemorrhage in CT images using encoder-decoder convolutional neural network
Kai Hu 0002, Kai Chen 0027, Xizhi He, Yuan Zhang 0022, Zhineng Chen, Xuanya Li, Xieping Gao 0001 |
Inf. Process. Manag. | 2 |