Zheng Yuan 0005

dblp:56/2877-5 · DBLP profile ↗
← Back
11ranked-venue papers
7as first author
10since 2021 · last 2026
0000-0001-8788-6817ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 5 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 5 first-author · 6 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 VLBiasBench: A Comprehensive Benchmark for Evaluating Bias in Large Vision-Language Model
abstract
The emergence of Large Vision-Language Models (LVLMs) marks significant strides towards achieving general artificial intelligence. However, these advancements are accompanied by concerns about biased outputs, a challenge that has yet to be thoroughly explored. Existing benchmarks are not sufficiently comprehensive in evaluating biases due to their limited data scale, single questioning format and narrow sources of bias. To address this problem, we introduce VLBiasBench, a comprehensive benchmark designed to evaluate biases in LVLMs. VLBiasBench features a dataset that covers nine distinct categories of social biases, including age, disability status, gender, nationality, physical appearance, race, religion, profession, social economic status, as well as two intersectional bias categories: race × gender and race × social economic status. To build a large-scale dataset, we use Stable Diffusion XL model to generate 46,848 high-quality images, which are combined with various questions to create 128,342 samples. These questions are divided into open-ended and close-ended types, ensuring thorough consideration of bias sources and a comprehensive evaluation of LVLM biases from multiple perspectives. We conduct extensive evaluations on 15 open-source models as well as two advanced closed-source models, yielding new insights into the biases present in these models.
Sibo Wang 0012, Xiangkui Cao, Jie Zhang 0071, Zheng Yuan 0005, Shiguang Shan, Xilin Chen 0001, Wen Gao 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 Dysca: A Dynamic and Scalable Benchmark for Evaluating Perception Ability of LVLMs
abstract
Currently many benchmarks have been proposed to evaluate the perception ability of the Large Vision-Language Models (LVLMs). However, most benchmarks conduct questions by selecting images from existing datasets, resulting in the potential data leakage. Besides, these benchmarks merely focus on evaluating LVLMs on the realistic style images and clean scenarios, leaving the multi-stylized images and noisy scenarios unexplored. In response to these challenges, we propose a dynamic and scalable benchmark named Dysca for evaluating LVLMs by leveraging synthesis images. Specifically, we leverage Stable Diffusion and design a rule-based method to dynamically generate novel images, questions and the corresponding answers. We consider 51 kinds of image styles and evaluate the perception capability in 20 subtasks. Moreover, we conduct evaluations under 4 scenarios (i.e., Clean, Corruption, Print Attacking and Adversarial Attacking) and 3 question types (i.e., Multi-choices, True-or-false and Free-form). Thanks to the generative paradigm, Dysca serves as a scalable benchmark for easily adding new subtasks and scenarios. A total of 24 advanced open-source LVLMs and 2 close-source LVLMs are evaluated on Dysca, revealing the drawbacks of current LVLMs. The benchmark is released in anonymous github page \url{https://github.com/Benchmark-Dysca/Dysca}.
Jie Zhang 0071, Mengqi Lei, Zheng Yuan 0005, Bei Yan, Shiguang Shan, Xilin Chen 0001
ICLR4
2025 FullLoRA: Efficiently Boosting the Robustness of Pretrained Vision Transformers
abstract
In recent years, the Vision Transformer (ViT) model has gradually become mainstream in various computer vision tasks, and the robustness of the model has received increasing attention. However, existing large models tend to prioritize performance during training, potentially neglecting the robustness, which may lead to serious security concerns. In this paper, we establish a new challenge: exploring how to use a small number of additional parameters for adversarial finetuning to quickly and effectively enhance the adversarial robustness of a standardly trained model. To address this challenge, we develop novel LNLoRA module, incorporating a learnable layer normalization before the conventional LoRA module, which helps mitigate magnitude differences in parameters between the adversarial and standard training paradigms. Furthermore, we propose the FullLoRA framework by integrating the learnable LNLoRA modules into all key components of ViT-based models while keeping the pretrained model frozen, which can significantly improve the model robustness via adversarial finetuning in a parameter-efficient manner. Extensive experiments on several datasets demonstrate the superiority of our proposed FullLoRA framework. It achieves comparable robustness with full finetuning while only requiring about 5% of the learnable parameters. This also effectively addresses concerns regarding extra model storage space and enormous training time caused by adversarial finetuning.
Zheng Yuan 0005, Jie Zhang 0071, Shiguang Shan, Xilin Chen 0001
IEEE Trans. Image Process.1
2024 Pre-Trained Model Guided Fine-Tuning for Zero-Shot Adversarial Robustness
abstract
Large-scale pre-trained vision-language models like CLIP have demonstrated impressive performance across various tasks, and exhibit remarkable zero-shot generalization capability, while they are also vulnerable to impercep-tible adversarial examples. Existing works typically em-ploy adversarial training (fine-tuning) as a defense method against adversarial examples. However, direct application to the CLIP model may result in overfitting, compromising the model's capacity for generalization. In this paper, we propose Pre-trained Model Guided Adversarial Fine-Tuning (PMG-AFT) method, which leverages supervision from the original pre-trained model by carefully designing an auxiliary branch, to enhance the model's zero-shot ad-versarial robustness. Specifically, PMG-AFT minimizes the distance between the features of adversarial examples in the target model and those in the pre-trained model, aiming to preserve the generalization features already captured by the pre-trained model. Extensive Experiments on 15 zero-shot datasets demonstrate that PMG-AFT significantly outper-forms the state-of-the-art method, improving the top-1 ro-bust accuracy by an average of 4.99%. Furthermore, our approach consistently improves clean accuracy by an aver-age of 8.72%. Our code is available at here.1
Sibo Wang 0012, Jie Zhang 0071, Zheng Yuan 0005, Shiguang Shan
CVPR3
2024 Towards Robust Semantic Segmentation against Patch-Based Attack via Attention Refinement
Zheng Yuan 0005, Jie Zhang 0071, Yude Wang, Shiguang Shan, Xilin Chen 0001
Int. J. Comput. Vis.1
2024 Adaptive Perturbation for Adversarial Attack
abstract
In recent years, the security of deep learning models achieves more and more attentions with the rapid development of neural networks, which are vulnerable to adversarial examples. Almost all existing gradient-based attack methods use the sign function in the generation to meet the requirement of perturbation budget on$L_\infty$norm. However, we find that the sign function may be improper for generating adversarial examples since it modifies the exact gradient direction. Instead of using the sign function, we propose to directly utilize the exact gradient direction with a scaling factor for generating adversarial perturbations, which improves the attack success rates of adversarial examples even with fewer perturbations. At the same time, we also theoretically prove that this method can achieve better black-box transferability. Moreover, considering that the best scaling factor varies across different images, we propose an adaptive scaling factor generator to seek an appropriate scaling factor for each image, which avoids the computational cost for manually searching the scaling factor. Our method can be integrated with almost all existing gradient-based attack methods to further improve their attack success rates. Extensive experiments on the CIFAR10 and ImageNet datasets show that our method exhibits higher transferability and outperforms the state-of-the-art methods.
Zheng Yuan 0005, Jie Zhang 0071, Zhaoyan Jiang, Shiguang Shan
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 Adaptive Adversarial Patch Attack on Face Recognition Models
abstract
Face recognition models have become widely used for identity authentication in scenarios such as cell phone unlocking and financial payment, but they are vulnerable to adversarial examples. Due to the realizability in the physical world, adversarial patch attack has emerged as a significant security threat. However, most existing adversarial patch attack methods focus on only one aspect of patch generation, such as patch location or shape. To overcome this limitation, we propose a novel unified Adaptive Adversarial Patch (AAP) attack framework for targeted attack on face recognition models. Our method comprehensively considers various factors during patch generation, including location, shape, and number. Our approach adaptively selects patch location and number based on saliency map and clustering, while simultaneously deforming patch shape and optimizing perturbations. Extensive experiments under both white-box and black-box settings demonstrate that our proposed method achieves higher attack success rates compared to SOTA methods.
Bei Yan, Jie Zhang 0071, Zheng Yuan 0005, Shiguang Shan
IJCB3
2023 RAMM: Retrieval-augmented Biomedical Visual Question Answering with Multi-modal Pre-training
abstract
Vision-and-language multi-modal pretraining and fine-tuning have shown great success in visual question answering (VQA). Compared to general domain VQA, the performance of biomedical VQA suffers from limited data. In this paper, we propose a retrieval-augmented pretrain-and-finetune paradigm named RAMM for biomedical VQA to overcome the data limitation issue. Specifically, we collect a new biomedical dataset named PMCPM which offers patient-based image-text pairs containing diverse patient situations from PubMed. Then, we pretrain the biomedical multi-modal model to learn visual and textual representation for image-text pairs and align these representations with image-text contrastive objective (ITC). Finally, we propose a retrieval-augmented method to better use the limited data. We propose to retrieve similar image-text pairs based on ITC from pretraining datasets and introduce a novel retrieval-attention module to fuse the representation of the image and the question with the retrieved images and texts. Experiments demonstrate that our retrieval-augmented pretrain-and-finetune paradigm obtains state-of-the-art performance on Med-VQA2019, Med-VQA2021, VQARAD, and SLAKE datasets. Further analysis shows that the proposed RAMM and PMCPM can enhance biomedical VQA performance compared with previous resources and methods. The pre-trained models and codes are published at https://github.com/GanjinZero/RAMM.
Zheng Yuan 0005, Qiao Jin 0001, Chuanqi Tan, Zhengyun Zhao, Hongyi Yuan, Fei Huang 0002, Songfang Huang
ACM Multimedia1
2022 Adaptive Image Transformations for Transfer-Based Adversarial Attack
Zheng Yuan 0005, Jie Zhang 0071, Shiguang Shan
ECCV (5)1
2021 Meta Gradient Adversarial Attack
abstract
In recent years, research on adversarial attacks has be-come a hot spot. Although current literature on the transfer-based adversarial attack has achieved promising results for improving the transferability to unseen black-box models, it still leaves a long way to go. Inspired by the idea of meta-learning, this paper proposes a novel architecture called Meta Gradient Adversarial Attack (MGAA), which is plug-and-play and can be integrated with any existing gradient-based attack method for improving the cross-model transferability. Specifically, we randomly sample multiple models from a model zoo to compose different tasks and iteratively simulate a white-box attack and a black-box attack in each task. By narrowing the gap between the gradient directions in white-box and black-box attacks, the transfer-ability of adversarial examples on the black-box setting can be improved. Extensive experiments on the CIFAR10 and ImageNet datasets show that our architecture outperforms the state-of-the-art methods for both black-box and white-box attack settings.
Zheng Yuan 0005, Jie Zhang 0071, Yunpei Jia, Chuanqi Tan, Shiguang Shan
ICCV1
2020 Attributes Aware Face Generation with Generative Adversarial Networks
abstract
Recent studies have shown remarkable success in face image generations. However, most of the existing methods only generate face images from random noise, and cannot generate face images according to the specific attributes. In this paper, we focus on the problem of face synthesis from attributes, which aims at generating faces with specific characteristics corresponding to the given attributes. To this end, we propose a novel attributes aware face image generator method with generative adversarial networks called AFGAN. Specifically, we firstly propose a two-path embedding layer and self-attention mechanism to convert binary attribute vector to rich attribute features. Then three stacked generators generate 64 × 64, 128 × 128 and 256 × 256 resolution face images respectively by taking the attribute features as input. In addition, an image-attribute matching loss is proposed to enhance the correlation between the generated images and input attributes. Extensive experiments on CelebA demonstrate the superiority of our AFGAN in terms of both qualitative and quantitative evaluations.
Zheng Yuan 0005, Jie Zhang 0071, Shiguang Shan, Xilin Chen 0001
ICPR1