Wenxuan Wang 0003

dblp:203/1536-3 · DBLP profile ↗
← Back
18ranked-venue papers
7as first author
14since 2021 · last 2026
0000-0001-7142-9734ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 6 first-author · 9 since 2021Artificial intelligence and machine learning · 12 · 4 first-author · 9 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Post-enhancing viewpoint robustness for pre-training generalizable re-identification
Yunlong Wang 0008, Bingliang Jiao, Liying Gao, Wenxuan Wang 0003, Peng Wang 0015
Knowl. Based Syst.4
2026 Lifelong person re-identification via dynamically knowledge adaptation and retention
Bingliang Jiao, Wenxuan Wang 0003, Peng Wang 0015
Neural Networks3
2026 Words Strip Wardrobe: Learning Invariant Textual Prompts for Cloth-Changing Person Re-Identification
abstract
Cloth-changing person re-identification (CC-ReID) aims to match individuals wearing varying clothes across camera views. Existing CC-ReID methods typically focus on extracting clothing-invariant features such as body shape, pose, gait,etc. However, these features are diverse and often entangled with clothing-related visual clues, posing significant challenges for comprehensively and effectively separating and leveraging them for re-identification. To address these challenges, we propose a text-guided clothing generalizable (Tex-CG) model, which employs multi-modal large language models (MLLMs) to comprehensively and explicitly decouple clothing-invariant features from pedestrian images in the textual domain. By shifting feature disentanglement to the textual domain, the interference caused by visual entanglement between clothing and clothing-invariant clues can be significantly reduced. Additionally, to ensure compatibility between offline-decoupled features from MLLMs and our online-trained Tex-CG model, we utilize CLIP-based image-text matching to train implicit clothing-invariant prompts that embed discriminative pedestrian information. A dynamic fusion module is subsequently introduced to leverage these implicit prompts for selectively integrating valuable and compatible components from the MLLM’s explicitly decoupled features, constructing robust guidance to direct our model to effectively capture clothing-invariant visual clues for re-identification. Extensive experiments demonstrate the effectiveness of our method, and the Tex-CG model achieves state-of-the-art performance on five mainstream CC-ReID benchmarks. Our code is available at https://github.com/JiaoBL1234/Tex-CG.
Bingliang Jiao, Liying Gao, Feiyue Zhao, Dapeng Oliver Wu, Wenxuan Wang 0003, Peng Wang 0015
IEEE Trans. Inf. Forensics Secur.5
2025 SeCap: Self-Calibrating and Adaptive Prompts for Cross-view Person Re-Identification in Aerial-Ground Networks
abstract
When discussing the Aerial-Ground Person Re-Identification (AGPReID) task, we face the main challenge of the significant appearance variations caused by different viewpoints, making identity matching difficult. To address this issue, previous methods attempt to reduce the differences between viewpoints by critical attributes and decoupling the viewpoints. While these methods can mitigate viewpoint differences to some extent, they still face two main issues: (1) difficulty in handling viewpoint diversity and (2) neglect of the contribution of local features. To effectively address these challenges, we design and implement the Self-Calibrating and Adaptive Prompt (SeCap) method for the AGPReID task. The core of this framework relies on the Prompt Re-Calibration Module (PRM), which adaptively re-calibrates prompts based on the input. Combined with the Local Feature Refinement Module (LFRM), SeCap can extract view-invariant features from local features for AGPReID. Meanwhile, given the current scarcity of datasets in the AGPReID field, we further contribute two real-world Large-Scale Aerial-Ground Person Re-Identification datasets, LAGPeR and G2APS-ReID. The former is collected and annotated by us independently, covering 4, 231 unique identities and containing 63, 841 high-quality images; the latter is reconstructed from the person search dataset G2APS. Through extensive experiments on AGPReID datasets, we demonstrate that SeCap is a feasible and effective solution for the AGPReID task. The datasets and source code available on https://github.com/wangshining681/SeCap-AgpreID.
Shining Wang, Yunlong Wang 0008, Ruiqi Wu 0001, Bingliang Jiao, Wenxuan Wang 0003, Peng Wang 0015
CVPR5
2025 Improving the Transferability of Adversarial Attacks on Face Recognition with Diverse Parameters Augmentation
abstract
Face Recognition (FR) models are vulnerable to adversarial examples that subtly manipulate benign face images, underscoring the urgent need to improve the transferability of adversarial attacks in order to expose the blind spots of these systems. Existing adversarial attack methods often overlook the potential benefits of augmenting the surrogate model with diverse initializations, which limits the transferability of the generated adversarial examples. To address this gap, we propose a novel method called Diverse Parameters Augmentation (DPA) attack method, which enhances surrogate models by incorporating diverse parameter initializations, resulting in a broader and more diverse set of surrogate models. Specifically, DPA consists of two key stages: Diverse Parameters Optimization (DPO) and Hard Model Aggregation (HMA). In the DPO stage, we initialize the parameters of the surrogate model using both pre-trained and random parameters. Subsequently, we save the models in the intermediate training process to obtain a diverse set of surrogate models. During the HMA stage, we enhance the feature maps of the diversified surrogate models by incorporating beneficial perturbations, thereby further improving the transferability. Experimental results demonstrate that our proposed attack method can effectively enhance the transferability of the crafted adversarial face examples.
Fengfan Zhou, Bangjie Yin, Qianyu Zhou 0001, Wenxuan Wang 0003
CVPR5
2025 Hierarchical Context Interaction and Reasoning with Transformer for Emotion Recognition
abstract
Emotion recognition is an important task in computer vision. However, current approaches using hard associations (e.g., element-wise addition or concatenation) suffer from information pollution. To overcome these challenges and utilize information at different scales, we present a novel Transformer-based emotion recognition (TransEmo) framework utilizing only the RGB modality as input. TransEmo is a hierarchical feature fusion framework with soft-association to effectively realize contextual interaction and emotion reasoning. It comprises a top-down path and a bottom-up path with cross-attention layers. The top-down path enriches the features of the current hierarchy gradually by using information from a larger receptive field and interacts effectively with contexts to suppress redundant information and noise. The bottom-up path introduces a set of learnable prototypes to adaptively capture clues for various emotions by reasoning with hierarchical features. Our proposed approach achieves state-of-the-art performance on two benchmarks and provides an analysis to discuss the advantages of our modules.
Wenxuan Wang 0003, Chenglei Wang, Xuli Shen, Qing Xu 0017, Xuelin Qian
ICASSP1
2025 Latent-space diffusion models for stealthy and transferable adversarial attacks on object detection
Wenxuan Wang 0003, Huihui Qi, Zhixiang Huang, Bangjie Yin, Peng Wang 0015
Neurocomputing1
2025 Dynamic Routing and Knowledge Re-Learning for Data-Free Black-Box Attack
abstract
Deep learning models have emerged as strong and efficient tools that can be applied to a broad spectrum of complex learning problems and many real-world applications. However, more and more works show that deep models are vulnerable to adversarial examples. Compared to vanilla attack settings, this paper advocates a more practical setting of data-free black-box attack, for which the attackers can completely not access the structures and parameters of the target model, as well as the intermediate features and any training data associated with the model. To tackle this task, previous methods generate transferable adversarial examples from a transparent substitute model to the target model. However, we found that these works have the limitations of taking static substitute model structure for different targets, only using hard synthesized examples once, and still relying on data statistics of the target model. This may potentially harm the performance of attacking the target model. To this end, we propose a novel Dynamic Routing and Knowledge Re-Learning framework (DraKe) to effectively learn a dynamic substitute model from the target model. Specifically, given synthesized training samples, a dynamic substitute structure learning strategy is proposed to adaptively generate optimal substitute model structure via a policy network according to different target models and tasks. To facilitate the substitute training, we present a graph-based structure information learning to capture the structural knowledge learned from the target model. For the inherent limitation that online data generation can only be learned once, a dynamic knowledge re-learning strategy is proposed to adjust the weights of optimization objectives and re-learn hard samples. Extensive experiments on four public image classification datasets and one face recognition benchmark are conducted to evaluate the efficacy of our Drake. We can obtain significant improvement compared with state-of-the-art competitors. More importantly, our DraKe consistently achieves attack superiority for different target models (e.g., residual networks, and vision transformers), showing great potential for complex real-world applications.
Xuelin Qian, Wenxuan Wang 0003, Yu-Gang Jiang 0001, Xiangyang Xue 0001, Yanwei Fu 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 Enhancing Visible-Infrared Person Re-Identification With Modality- and Instance-Aware Adaptation Learning
abstract
The Visible-Infrared Person Re-identification (VI ReID) aims to achieve cross-modality re-identification by matching pedestrian images from visible and infrared illumination. A crucial challenge in this task is mitigating the impact of modality divergence to enable the VI ReID model to learn cross-modality correspondence. Regarding this challenge, existing methods primarily focus on eliminating the information gap between different modalities by extracting modality-invariant information or supplementing inputs with specific information from another modality. However, these methods may overly focus on bridging the information gap, a challenging issue that could potentially overshadow the inherent complexities of cross-modality ReID itself. Based on this insight, we propose a straightforward yet effective strategy to empower the VI ReID model with sufficient flexibility to adapt diverse modality inputs to achieve cross-modality ReID effectively. Specifically, we introduce a Modality-aware and Instance-aware Visual Prompts (MIP) network, leveraging transformer architecture with customized visual prompts. In our MIP, a set of modality-aware prompts is designed to enable our model to dynamically adapt diverse modality inputs and effectively extract information for identification, thereby alleviating the interference of modality divergence. Besides, we also propose the instance-aware prompts, which are responsible for guiding the model to adapt individual pedestrians and capture discriminative clues for accurate identification. Through extensive experiments on four mainstream VI ReID datasets, the effectiveness of our designed modules is evaluated. Furthermore, our proposed MIP network outperforms most current state-of-the-art methods.
Ruiqi Wu 0001, Bingliang Jiao, Shining Wang, Wenxuan Wang 0003, Peng Wang 0015
IEEE Trans. Circuits Syst. Video Technol.5
2024 Enhancing Visible-Infrared Person Re-identification with Modality- and Instance-aware Visual Prompt Learning
abstract
The Visible-Infrared Person Re-identification (VI ReID) aims to match visible and infrared images of the same pedestrians across non-overlapped camera views. These two input modalities contain both invariant information, such as shape, and modality-specific details, such as color. An ideal model should utilize valuable information from both modalities during training for enhanced representational capability. However, the gap caused by modality-specific information poses substantial challenges for the VI ReID model to handle distinct modality inputs simultaneously. To address this, we introduce the Modality-aware and Instance-aware Visual Prompts (MIP) network in our work, designed to effectively utilize both invariant and specific information for identification. Specifically, our MIP model is built on the transformer architecture. In this model, we have designed a series of modality-specific prompts, which could enable our model to adapt to and make use of the specific information inherent in different modality inputs, thereby reducing the interference caused by the modality gap and achieving better identification. Besides, we also employ each pedestrian feature to construct a group of instance-specific prompts. These customized prompts are responsible for guiding our model to adapt to each pedestrian instance dynamically, thereby capturing identity-level discriminative clues for identification. Through extensive experiments on SYSU-MM01 and RegDB datasets, the effectiveness of both our designed modules is evaluated. Additionally, our proposed MIP performs better than most state-of-the-art methods.
Ruiqi Wu 0001, Bingliang Jiao, Wenxuan Wang 0003, Peng Wang 0015
ICMR3
2024 Sustainable Self-evolution Adversarial Training
abstract
With the wide application of deep neural network models in various computer vision tasks, there has been a proliferation of adversarial example generation strategies aimed at deeply exploring model security. However, existing adversarial training defense models, which rely on single or limited types of attacks under a one-time learning process, struggle to adapt to the dynamic and evolving nature of attack methods. Therefore, to achieve defense performance improvements for models in long-term applications, we propose a novel Sustainable Self-Evolution Adversarial Training (SSEAT) framework. Specifically, we introduce a continual adversarial defense pipeline to realize learning from various kinds of adversarial examples across multiple stages. Additionally, to address the issue of model catastrophic forgetting caused by continual learning from ongoing novel attacks, we propose an adversarial data replay module to better select more diverse and key relearning data. Furthermore, we design a consistency regularization strategy to encourage current defense models to learn more from previously trained ones, guiding them to retain more past knowledge and maintain accuracy on clean samples. Extensive experiments have been conducted to verify the efficacy of the proposed SSEAT defense method, which demonstrates superior defense performance and classification accuracy compared to competitors.
Wenxuan Wang 0003, Chenglei Wang, Huihui Qi, Menghao Ye, Xuelin Qian, Peng Wang 0015, Yanning Zhang 0001
ACM Multimedia1
2022 DST: Dynamic Substitute Training for Data-free Black-box Attack
abstract
With the wide applications of deep neural network models in various computer vision tasks, more and more works study the model vulnerability to adversarial examples. For data-free black box attack scenario, existing methods are inspired by the knowledge distillation, and thus usually train a substitute model to learn knowledge from the target model using generated data as input. However, the substitute model always has a static network structure, which limits the attack ability for various target models and tasks. In this paper, we propose a novel dynamic substitute training attack method to encourage substitute model to learn better and faster from the target model. Specifically, a dynamic substitute structure learning strategy is proposed to adaptively generate optimal substitute model structure via a dy-namic gate according to different target models and tasks. Moreover, we introduce a task-driven graph-based structure information learning constrain to improve the quality of generated training data, and facilitate the substitute model learning structural relationships from the target model multiple outputs. Extensive experiments have been conducted to verify the efficacy of the proposed attack method, which can achieve better performance compared with the state-of-the-art competitors on several datasets. Project page: https://wxwangiris.github.io/DST
Wenxuan Wang 0003, Xuelin Qian, Yanwei Fu 0001, Xiangyang Xue 0001
CVPR1
2021 Delving into Data: Effectively Substitute Training for Black-box Attack
abstract
Deep models have shown their vulnerability when processing adversarial samples. As for the black-box attack, without access to the architecture and weights of the attacked model, training a substitute model for adversarial attacks has attracted wide attention. Previous substitute training approaches focus on stealing the knowledge of the target model based on real training data or synthetic data, without exploring what kind of data can further improve the transferability between the substitute and target models. In this paper, we propose a novel perspective substitute training that focuses on designing the distribution of data used in the knowledge stealing process. More specifically, a diverse data generation module is proposed to synthesize large-scale data with wide distribution. And adversarial substitute training strategy is introduced to focus on the data distributed near the decision boundary. The combination of these two modules can further boost the consistency of the substitute model and target model, which greatly improves the effectiveness of adversarial attack. Extensive experiments demonstrate the efficacy of our method against state-of-the-art competitors under non-target and target at-tack settings. Detailed visualization and analysis are also provided to help understand the advantage of our method.
Wenxuan Wang 0003, Bangjie Yin, Taiping Yao, Li Zhang 0040, Yanwei Fu 0001, Shouhong Ding, Feiyue Huang, Xiangyang Xue 0001
CVPR1
2021 Adv-Makeup: A New Imperceptible and Transferable Attack on Face Recognition
abstract
Deep neural networks, particularly face recognition models, have been shown to be vulnerable to both digital and physical adversarial examples. However, existing adversarial examples against face recognition systems either lack transferability to black-box models, or fail to be implemented in practice. In this paper, we propose a unified adversarial face generation method - Adv-Makeup, which can realize imperceptible and transferable attack under the black-box setting. Adv-Makeup develops a task-driven makeup generation method with the blending module to synthesize imperceptible eye shadow over the orbital region on faces. And to achieve transferability, Adv-Makeup implements a fine-grained meta-learning based adversarial attack strategy to learn more vulnerable or sensitive features from various models. Compared to existing techniques, sufficient visualization results demonstrate that Adv-Makeup is capable to generate much more imperceptible attacks under both digital and physical scenarios. Meanwhile, extensive quantitative experiments show that Adv-Makeup can significantly improve the attack success rate under black-box setting, even attacking commercial systems.
Bangjie Yin, Wenxuan Wang 0003, Taiping Yao, Zelun Kong, Shouhong Ding, Cong Liu 0005
IJCAI2
2020 Long-Term Cloth-Changing Person Re-identification
Xuelin Qian, Wenxuan Wang 0003, Li Zhang 0040, Fangrui Zhu, Yanwei Fu 0001, Tao Xiang 0002, Yu-Gang Jiang 0001, Xiangyang Xue 0001
ACCV (3)2
2020 FM2u-Net: Face Morphological Multi-Branch Network for Makeup-Invariant Face Verification
abstract
It is challenging in learning a makeup-invariant face verification model, due to (1) insufficient makeup/non-makeup face training pairs, (2) the lack of diverse makeup faces, and (3) the significant appearance changes caused by cosmetics. To address these challenges, we propose a unified Face Morphological Multi-branch Network (FMMu-Net) for makeup-invariant face verification, which can simultaneously synthesize many diverse makeup faces through face morphology network (FM-Net) and effectively learn cosmetics-robust face representations using attention-based multi-branch learning network (AttM-Net). For challenges (1) and (2), FM-Net (two stacked auto-encoders) can synthesize realistic makeup face images by transferring specific regions of cosmetics via cycle consistent loss. For challenge (3), AttM-Net, consisting of one global and three local (task-driven on two eyes and mouth) branches, can effectively capture the complementary holistic and detailed information. Unlike DeepID2 which uses simple concatenation fusion, we introduce a heuristic method AttM-FM, attached to AttM-Net, to adaptively weight the features of different branches guided by the holistic information. We conduct extensive experiments on makeup face verification benchmarks (M-501, M-203, and FAM) and general face recognition datasets (LFW and IJB-A). Our framework FMMu-Net achieves state-of-the-art performances.
Wenxuan Wang 0003, Yanwei Fu 0001, Xuelin Qian, Yu-Gang Jiang 0001, Qi Tian 0001, Xiangyang Xue 0001
CVPR1
2019 Comp-GAN: Compositional Generative Adversarial Network in Synthesizing and Recognizing Facial Expression
abstract
Facial expression is important in understanding our social interaction. Thus the ability to recognize facial expression enables the novel multimedia applications. With the advance of recent deep architectures, research on facial expression recognition has achieved great progress. However, these models are still suffering from the problems of lacking sufficient and diverse high quality training faces, vulnerability to the facial variations, and recognizing a limited number of basic types of emotions. To tackle these problems, this paper proposes a novel end-to-end Compositional Generative Adversarial Network (Comp-GAN) that is able to synthesize new face images with specified poses and desired facial expressions; and such synthesized images can be further utilized to help train a robust and generalized expression recognition model. Essentially, Comp-GAN can dynamically change the expression and pose of faces according to the input images while keeping the identity information. Specifically, the generator has two major components: one for generating images with desired expression and the other for changing the pose of faces. Furthermore, a face reconstruction learning process is applied to re-generate the input image and constrains the generator for preserving the key information such as facial identity. For the first time, various one/zero-shot facial expression recognition tasks have been created. We conduct extensive experiments to show that the images generated by Comp-GAN are helpful to improve the performance of one/zero-shot facial expression recognition.
Wenxuan Wang 0003, Qiang Sun 0007, Yanwei Fu 0001, Tao Chen 0003, Chenjie Cao, Ziqi Zheng, Han Qiu 0002, Yu-Gang Jiang 0001, Xiangyang Xue 0001
ACM Multimedia1
2018 Pose-Normalized Image Generation for Person Re-identification
Xuelin Qian, Yanwei Fu 0001, Tao Xiang 0002, Wenxuan Wang 0003, Yang Wu 0001, Yu-Gang Jiang 0001, Xiangyang Xue 0001
ECCV (9)4