Qiqi Bao 0001

dblp:215/7768-1 · DBLP profile ↗
← Back
19ranked-venue papers
5as first author
19since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 4 first-author · 12 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Security and privacy · 2 · 2 since 2021
YearPublicationVenuePosition
2026 EFSRNet: Multi-scale exposure normalization and dual-branch aggregation for overexposed face super-resolution
Shengying Yang, Bingxin Zha, Qiqi Bao 0001, Boyang Feng, Zhiwang Xu
Knowl. Based Syst.4
2026 RPA: Recursive Perturbation-Based Universal Adversarial Attacks on Multimodal Generative Tasks
abstract
Current adversarial attacks pose a serious threat to the robustness of visual-language models (VLMs), including vision-language pre-trained models (VLPMs) and multimodal large language models (MLLMs). Traditional adversarial attacks are example-specific and rely on specific datasets. This practice suffers from low transferability and additional computation cost, while universal adversarial perturbations (UAPs) offer example-agnostic solutions by generalizing across inputs. However, current UAP methods mainly target VLPMs, demonstrating limited transferability and effectiveness in MLLMs. To bridge this gap, we propose the Recursive Perturbation Attack (RPA), a novel black-box UAP method for both VLPMs and MLLMs. RPA employs a recursive perturbations strategy, utilizing token filtering and polynomial sampling methods to generate perturbations, thereby achieving incremental disruption and enhancing the transferability of the attack. To further enhance the effectiveness of the attack, RPA integrates a three-tier modality decoupling strategy, disentangling intra-modal, cross-modal, and fusion-modal features to effectively disrupt feature alignment and interactions. Extensive experiments validate that RPA achieves superior attack performance compared to existing UAP approaches. This work highlights new security concerns in multimodal AI systems and provides insights into the design of more robust models. Code is available at https://github.com/chilljudaoren/RPAttack.
Yaguan Qian, Qiqi Bao 0001, Chang Zong, Fei Yu 0012, Shouling Ji, Bin Wang 0062, Zhaoquan Gu, Zhen Lei 0001
IEEE Trans. Circuits Syst. Video Technol.3
2026 Exploiting Shared Adversarial Features for Dynamic Attacks in Large Vision-Language Models
abstract
With the rapid development of Large Language Models (LLMs), an increasing number of Large Visual-Language Models (LVLMs) have achieved unprecedented performance in response generation. Recent work shows that LVLMs are vulnerable to adversarial attacks. However, many existing methods tend to overfit to the source model by overemphasizing specific features, which compromises their transferability. Other approaches suffer from reduced attack effectiveness due to insufficient differentiation between features. In this paper, we propose a novel transfer-based black-box untargeted attack—Shared Adversarial Feature (SAF) dynamic attack. By exploring the feature extraction patterns of LVLMs, we identify the features shared among various models that are most susceptible to adversarial attacks and disrupt them. Moreover, due to the powerful attention mechanisms of LVLMs, they are still able to extract similar semantics from perturbed images, even when primary features are disrupted. We design a dynamic update strategy to address this challenge. Finally, from the perspective of SAF, we conduct an in-depth analysis of vulnerabilities in the vision encoder and projector within LVLMs and find that attacking the projector exhibits stronger transferability across heterogeneous model architectures. Extensive experiments show that our method exhibits superior attack performance compared to existing methods across different models, datasets, and tasks. The code will be publicly available after publication.
Yaguan Qian, Xucheng Zhu, Qiqi Bao 0001, Fei Yu 0012, Shouling Ji, Zhaoquan Gu, Wei Wang 0012, Bin Wang 0062, Zhen Lei 0001
IEEE Trans. Inf. Forensics Secur.3
2026 Individual and Common Attack: Enhancing Transferability in VLP Models Through Modal Feature Exploitation
abstract
Vision-Language Pretrained (VLP) models exhibit strong multimodal understanding and reasoning capabilities, finding wide application in tasks such as image-text retrieval and visual grounding. However, they remain highly vulnerable to adversarial attacks, posing serious reliability concerns in safety-critical scenarios. We observe that existing adversarial examples optimization methods typically rely on individual features from the other modality as guidance, causing the crafted adversarial examples to overfit that modality's learning preferences and thus limiting their transferability. In order to further enhance the transferability of adversarial examples, we propose a novel adversarial attack framework, I&CA (Individual & Common feature Attack), which simultaneously considers individual features within each modality and common features cross-modal interactions. Concretely, I&CA first drives divergence among individual features within each modality to disrupt single-modality learning, and then suppresses the expression of common features during cross-modal interactions, thereby undermining the robustness of the fusion mechanism. In addition, to prevent adversarial perturbations from overfitting to the learning bias of the other modality, which may distort the representation of common features, we simultaneously introduce augmentation strategies to both modalities. Across various experimental settings and widely recognized multimodal benchmarks, the I&CA framework achieves an average transferability improvement of 6.15% over the state-of-the-art DRA method, delivering significant performance gains in both cross-model and cross-task attack scenarios.
Yaguan Qian, Yaxin Kong, Qiqi Bao 0001, Zhaoquan Gu, Bin Wang 0062, Shouling Ji, Zhen Lei 0001
IEEE Trans. Image Process.3
2025 Pose Magic: Efficient and Temporally Consistent Human Pose Estimation with a Hybrid Mamba-GCN Network
abstract
Current state-of-the-art (SOTA) methods in 3D Human Pose Estimation (HPE) are primarily based on Transformers. However, existing Transformer-based 3D HPE backbones often encounter a trade-off between accuracy and computational efficiency. To resolve the above dilemma, in this work, we leverage recent advances in state space models and utilize Mamba for high-quality and efficient long-range modeling. Nonetheless, Mamba still faces challenges in precisely exploiting local dependencies between joints. To address these issues, we propose a new attention-free hybrid spatiotemporal architecture named Hybrid Mamba-GCN (Pose Magic). This architecture introduces local enhancement with GCN by capturing relationships between neighboring joints, thus producing new representations to complement Mamba's outputs. By adaptively fusing representations from Mamba and GCN, Pose Magic demonstrates superior capability in learning the underlying 3D structure. To meet the requirements of real-time inference, we also provide a fully causal version. Extensive experiments show that Pose Magic achieves new SOTA results (0.9 mm drop) while saving 74.1% FLOPs. In addition, Pose Magic exhibits optimal motion consistency and the ability to generalize to unseen sequence lengths.
Xinyi Zhang 0008, Qiqi Bao 0001, Qinpeng Cui, Wenming Yang, Qingmin Liao
AAAI2
2025 Elucidating the Solution Space of Extended Reverse-Time SDE for Diffusion Models
abstract
Sampling from Diffusion Models can alternatively be seen as solving differential equations, where there is a challenge in balancing speed and image visual quality. ODE-based samplers offer rapid sampling time but reach a performance limit, whereas SDE-based samplers achieve superior quality, albeit with longer iterations. In this work, we formulate the sampling process as an Extended Reverse-Time SDE (ER SDE), unifying prior explorations into ODEs and SDEs. Theoretically, leveraging the semi-linear structure of ER SDE solutions, we offer exact solutions and approximate solutions for VP SDE and VE SDE, respectively. Based on the approximate solution space of the ER SDE, referred to as one-step prediction errors, we yield mathematical insights elucidating the rapid sampling capability of ODE solvers and the high-quality sampling ability of SDE solvers. Additionally, we unveil that VP SDE solvers stand on par with their VE SDE counterparts. Based on these findings, leveraging the dual advantages of ODE solvers and SDE solvers, we devise efficient high-quality samplers, namely ER-SDE-Solvers. Experimental results demonstrate that ER-SDE-Solvers achieve state-of-the-art performance across all stochastic samplers while maintaining efficiency of deterministic samplers. Specifically, on the ImageNet 128 × 128 dataset, ER-SDE-Solvers obtain 8.33 FID in only 20 function evaluations. Code is available at https://github.com/QinpengCui/ER-SDE-Solver
Qinpeng Cui, Xinyi Zhang 0008, Qiqi Bao 0001, Qingmin Liao
WACV3
2025 Enhancing robust generalization through appropriate adversarial example attack intensity
abstract
Deep Neural Networks (DNNs) are notoriously susceptible to adversarial examples. To mitigate the impact of well-designed adversarial attacks on network models, researchers have developed various defense mechanisms, among which adversarial training has emerged as one of the most effective strategies to date. Adversarial training aims to augment training data with adversarial examples, thus giving DNNs a certain degree of robustness to defend against adversarial attacks. However, while obtaining adversarial robustness, this method comes at the cost of reducing the generalization performance, manifested in the reduced classification effect of clean test datasets. Researchers have been actively seeking to counter the balance between adversarial robustness and model generalization. We believe that the key to balancing these two aspects lies in identifying appropriate adversarial examples. Overly potent examples can lead to a decline in clean accuracy, whereas weaker examples may offer limited robustness. Based on our analysis, a new adversarial example generation algorithm called Denoising Projection Gradient Descent (DPGD) was proposed. DPGD adds a purification module and a constraint in generating adversarial examples, the former is used to limit the influence of too strong adversarial examples on model training and the latter is used to ensure the necessary attack intensity. Combining DPGD with the framework of traditional adversarial training, we obtain the Diffusion Adversarial Training (DifAT) approach. To verify the effectiveness of our proposed method, we conducted extensive experiments on benchmark datasets, including CIFAR-10, CIFAR-100, and Tiny-Imagenet. Our results demonstrate the effectiveness of DifAT in improving the robustness of DNNs while maintaining or even improving their generalization performance.
Xiaoguo Ding, Liangjian Zhang, Qiqi Bao 0001, Yaguan Qian, Bin Wang 0062, Zhaoquan Gu, Yanchun Zhang
Neurocomputing3
2025 A Multimodal Adversarial Attack Method via Frequency Domain Enhancement and Fine-Grained Cross-Modal Guidance
abstract
Vision-language pretraining (VLP) models have demonstrated outstanding performance in image-text understanding tasks but remain highly susceptible to transferable adversarial attacks. While ensemble-based guided attacks improve adversarial transferability by increasing the diversity of image-text pairs, they primarily rely on spatial-domain data augmentation, which can lead to model overfitting to image details and limit the generalization capability of attacks. To address this limitation, this study proposes a frequency-domain adjustment-based adversarial attack method that modifies specific frequency components of input images to reduce detail interference and enhance the stability of adversarial examples. Additionally, a fine-grained feature extraction technique is introduced to optimize image-text alignment, further improving the transferability of cross-modal attacks. Experimental results demonstrate that the proposed method achieves superior attack transferability and generalization performance across two major VLP architectures, fusion models and alignment models, as well as multiple tasks on the Flickr30 K and MSCOCO datasets.
Yaguan Qian, Qinqin Yu, Qiqi Bao 0001, Shouling Ji, Wei Wang 0012, Bin Wang 0062, Zhaoquan Gu, Zhen Lei 0001
IEEE Trans. Dependable Secur. Comput.3
2024 Improving Diffusion-Based Image Restoration with Error Contraction and Error Correction
abstract
Generative diffusion prior captured from the off-the-shelf denoising diffusion generative model has recently attained significant interest. However, several attempts have been made to adopt diffusion models to noisy inverse problems either fail to achieve satisfactory results or require a few thousand iterations to achieve high-quality reconstructions. In this work, we propose a diffusion-based image restoration with error contraction and error correction (DiffECC) method. Two strategies are introduced to contract the restoration error in the posterior sampling process. First, we combine existing CNN-based approaches with diffusion models to ensure data consistency from the beginning. Second, to amplify the error contraction effects of the noise, a restart sampling algorithm is designed. In the error correction strategy, the estimation-correction idea is proposed on both the data term and the prior term. Solving them iteratively within the diffusion sampling framework leads to superior image generation results. Experimental results for image restoration tasks such as super-resolution (SR), Gaussian deblurring, and motion deblurring demonstrate that our approach can reconstruct high-quality images compared with state-of-the-art sampling-based diffusion models.
Qiqi Bao 0001, Zheng Hui, Rui Zhu 0006, Peiran Ren, Xuansong Xie, Wenming Yang
AAAI1
2024 Geometry-Guided Diffusion Model with Masked Transformer for Robust Multi-View 3D Human Pose Estimation
abstract
Recent research on Diffusion Models and Transformers has brought significant advancements to 3D Human Pose Estimation (HPE). Nonetheless, existing methods often fail to concurrently address the issues of accuracy and generalization. In this paper, we propose a Geometry-guided Dif fusion Model with Masked Transformer (Masked Gifformer) for robust multi-view 3D HPE. Within the framework of the diffusion model, a hierarchical multi-view trans-former-based denoiser is exploited to fit the 3D pose distribution by systematically integrating joint and view information. To address the long-standing problem of poor generalization, we introduce a fully random mask mechanism without any additional learnable modules or parameters. Furthermore, we incorporate geometric guidance into the diffusion model to enhance the accuracy of the model. This is achieved by optimizing the sampling process to minimize reprojection errors through modeling a conditional guidance distribution. Extensive experiments on two benchmarks demonstrate that Masked Gifformer effectively achieves a trade-off between accuracy and generalization. Specifically, our method outperforms other probabilistic methods by > 40% and achieves comparable results with state-of-the-art deterministic methods. In addition, our method exhibits robustness to varying camera numbers, spatial arrangements, and datasets.
Xinyi Zhang 0008, Qinpeng Cui, Qiqi Bao 0001, Wenming Yang, Qingmin Liao
ACM Multimedia3
2024 Taming Diffusion Prior for Image Super-Resolution with Domain Shift SDEs
abstract
Diffusion-based image super-resolution (SR) models have attracted substantial interest due to their powerful image restoration capabilities. However, prevailing diffusion models often struggle to strike an optimal balance between efficiency and performance. Typically, they either neglect to exploit the potential of existing extensive pretrained models, limiting their generative capacity, or they necessitate a dozens of forward passes starting from random noises, compromising inference efficiency. In this paper, we present DoSSR, a $\textbf{Do}$main $\textbf{S}$hift diffusion-based SR model that capitalizes on the generative powers of pretrained diffusion models while significantly enhancing efficiency by initiating the diffusion process with low-resolution (LR) images. At the core of our approach is a domain shift equation that integrates seamlessly with existing diffusion models. This integration not only improves the use of diffusion prior but also boosts inference efficiency. Moreover, we advance our method by transitioning the discrete shift process to a continuous formulation, termed as DoS-SDEs. This advancement leads to the fast and customized solvers that further enhance sampling efficiency. Empirical results demonstrate that our proposed method achieves state-of-the-art performance on synthetic and real-world datasets, while notably requiring $\textbf{\emph{only 5 sampling steps}}$. Compared to previous diffusion prior based methods, our approach achieves a remarkable speedup of 5-7 times, demonstrating its superior efficiency.
Qinpeng Cui, Yixuan Liu 0004, Xinyi Zhang 0008, Qiqi Bao 0001, Qingmin Liao, liwang Amd, Zicheng Liu 0001, Zhongdao Wang, Emad Barsoum
NeurIPS4
2023 EvenFace: Deep Face Recognition with Uniform Distribution of Identities
abstract
The development of loss functions over the past few years has brought great success to face recognition. Most algorithms focus on improving the intra-class compactness of face features but ignore the inter-class separability. In this paper, we propose a method named EvenFace, which introduces a regularization variance item and a mean term of inter-class separability to further promote the even distribution of class centers on the hypersphere, thereby increasing the inter-class distance. In order to evaluate the inter-class separability, a new index is proposed to better reflect the distribution of class centers and guide the classification. By penalizing the angle between each identity and its surrounding neighbors, the resulting uniform distribution of identities enables full exploitation of the feature space, leading to discriminative face representations. Our proposed loss function can effectively boost the performance of softmax loss variants. Quantitative comparisons with other state-of-the-art methods on several benchmarks demonstrate the superiority of EvenFace.
Yingfan Tao, Qiqi Bao 0001, Guijin Wang, Wenming Yang
ICME3
2023 SCTANet: A Spatial Attention-Guided CNN-Transformer Aggregation Network for Deep Face Image Super-Resolution
abstract
Numerous CNN-based algorithms have been proposed to reconstruct high-quality face images. However, the inability of convolution operation to model long-distance relationships limits the performance of the CNN-based methods. Moreover, in the high-resolution (HR) image reconstruction stage, with the well decoded feature representations, more efficient architecture design can be explored to synthesize pixel-level image details. In this work, we propose a spatial attention-guided CNN-Transformer aggregation network (SCTANet) for face image super-resolution (FSR) tasks. The core component in the deep feature extraction stage is the Hybrid Attention Aggregation (HAA) block. The HAA block has two parallel paths, one for the Residual Spatial Attention (RSA) block, the other for the Multi-scale Patch embedding and Spatial-attention Masked Transformer (MPSMT) block. The HAA block combines the strengths of CNN and transformer to effectively exploit both local and global information. For the reconstruction stage, we propose to use the Sub-pixel MLP-based Upsampling (SMU) module instead of the conventional CNN architecture. The SMU module promotes the reconstruction of pixel-level image details and reduces computational complexity. Extensive experiments on both synthetic and real-world face datasets demonstrate the superiority of our proposed SCTANet over state-of-the-art methods.
Qiqi Bao 0001, Yunmeng Liu, Bowen Gang, Wenming Yang, Qingmin Liao
IEEE Trans. Multim.1
2022 Quality-Oriented Feature Regression for Robust Image Similarity Metric
abstract
Full-reference image quality assessment aims to predict the perceptual quality of a distorted image based on its similarity to the pristine reference. In this paper, we propose a robust image similarity metric by fully exploring the representation power of deep learning-based features. A convolutional neu-ral network (CNN) is adopted to extract deep features from multiple scales. We show that such CNN features that con-tain multi -scale visual information are comprehensive and ro-bust enough for quality assessment. We further propose a quality-oriented feature regression (QOFR) module based on the multi-layer perceptron architecture. The QOFR module can efficiently integrate hierarchy CNN features and generate the final quality score. Extensive experiments on the bench-mark datasets demonstrate that our method achieves state-of-the-art performance with outstanding robustness and general-ization ability.
Qiqi Bao 0001, Rui Zhu 0006, Wenming Yang, Qingmin Liao
ICME2
2022 Distilling Resolution-robust Identity Knowledge for Texture-Enhanced Face Hallucination
abstract
The main focus of most existing face hallucination methods is to generate visually pleasing results. However, in many applications, the final goal is to identify the person in the low-resolution (LR) image. In this paper, we propose a texture and identity integration network (TIIN) to effectively incorporate identity information into face hallucination tasks. TIIN consists of an identity-preserving denormalization module (IDM) and an equalized texture enhance module (ETEM). The IDM exploits the identity prior and the ETEM improves image quality through histogram equalization. To extract identity information effectively, we propose a resolution-robust identity knowledge distillation network (RIKDN). RIKDN is specifically designed for LR face recognition and can be of independent interest. It employs two teacher-student streams. One stream narrows the performance gap between high-resolution (HR) and LR images. The other distills correlation information from the HR-HR teacher stream to guide learning in the LR-HR student stream. We conduct extensive experiments on multiple datasets to demonstrate the effectiveness of our methods.
Qiqi Bao 0001, Rui Zhu 0006, Bowen Gang, Pengyang Zhao, Wenming Yang, Qingmin Liao
ACM Multimedia1
2022 Attention-Driven Graph Neural Network for Deep Face Super-Resolution
abstract
With the help of convolutional neural networks (CNNs), deep learning-based methods have achieved remarkable performance in face super-resolution (FSR) task. Despite their success, most of the existing methods neglect non-local correlations of face images, leaving much room for improvement. In this paper, we introduce a novel end-to-end trainable attention-driven graph neural network (AD-GNN) for more discriminative feature extraction and feature relation modeling. This is achieved by two major components. The first component is a cross-scale dynamic graph (CDG) block. The CDG block considers cross-scale relationships of patches in distant areas and employs two dynamic graphs to construct enhanced features. The second component is a series of channel attention and spatial dynamic graph (CASDG) blocks. A CASDG block has a channel-wise attention unit and a spatial-aware dynamic graph (SDG) unit. The SDG unit extracts informative features by exploring spatial non-local self-similarity information of the patches using dynamic graph convolution. Using these two components, facial details can be effectively reconstructed with the help of information supplemented by similar but spatially remote patches and structural information of faces. Extensive experiments on two public benchmarks demonstrate the superiority of AD-GNN over the state-of-the-art FSR methods.
Qiqi Bao 0001, Bowen Gang, Wenming Yang, Jie Zhou 0001, Qingmin Liao
IEEE Trans. Image Process.1
2022 MDAN: Mirror Difference Aware Network for Brain Stroke Lesion Segmentation
abstract
Brain stroke lesion segmentation is of great importance for stroke rehabilitation neuroimaging analysis. Due to the large variance of stroke lesion shapes and similarities of tissue intensity distribution, it remains a challenging task. To help detect abnormalities, the anatomical symmetries of brain magnetic resonance (MR) images have been widely used as visual cues for clinical practices. However, most methods for brain images segmentation do not fully utilize structural symmetry information. This paper presents a novel mirror difference aware network (MDAN) for stroke lesion segmentation. The network uses an encoder-decoder architecture, aiming at holistically exploiting the symmetries of image features. Specifically, a differential feature augmentation (DFA) module is developed in the encoding path to highlight the semantically pathological asymmetries of features in abnormalities. In the DFA module, a Siamese contrastive supervised loss is designed to enhance discriminative features, and a mirror position-based difference augmentation (MDA) module is used to further magnify the discrepancy. Moreover, mirror feature fusion (MFF) modules are applied to efficiently fuse and transfer the information both of the original input and the horizontally flipped features to the decoding path. Extensive experiments on the Anatomical Tracings of Lesions After Stroke (ATLAS) dataset show the proposed MDAN outperforms the state-of-the-art methods.
Qiqi Bao 0001, Shiyu Mi, Bowen Gang, Wenming Yang, Jie Chen 0001, Qingmin Liao
IEEE J. Biomed. Health Informatics1
2022 RFormer: Transformer-Based Generative Adversarial Network for Real Fundus Image Restoration on a New Clinical Benchmark
abstract
Ophthalmologists have used fundus images to screen and diagnose eye diseases. However, different equipments and ophthalmologists pose large variations to the quality of fundus images. Low-quality (LQ) degraded fundus images easily lead to uncertainty in clinical screening and generally increase the risk of misdiagnosis. Thus, real fundus image restoration is worth studying. Unfortunately, real clinical benchmark has not been explored for this task so far. In this paper, we investigate the real clinical fundus image restoration problem. Firstly, We establish a clinical dataset, Real Fundus (RF), including 120 low- and high-quality (HQ) image pairs. Then we propose a novel Transformer-based Generative Adversarial Network (RFormer) to restore the real degradation of clinical fundus images. The key component in our network is the Window-based Self-Attention Block (WSAB) which captures non-local self-similarity and long-range dependencies. To produce more visually pleasant results, a Transformer-based discriminator is introduced. Extensive experiments on our clinical benchmark show that the proposed RFormer significantly outperforms the state-of-the-art (SOTA) methods. In addition, experiments of downstream tasks such as vessel segmentation and optic disc/cup detection demonstrate that our proposed RFormer benefits clinical fundus image analysis and applications.
Zhuo Deng 0001, Yuanhao Cai, Qiqi Bao 0001, Xue Yao, Wenming Yang, Shaochong Zhang
IEEE J. Biomed. Health Informatics5
2021 MBFF-Net: Multi-Branch Feature Fusion Network for Carotid Plaque Segmentation in Ultrasound
Shiyu Mi, Qiqi Bao 0001, Zhanghong Wei, Wenming Yang
MICCAI (5)2