Junyi Wu 0001

dblp:125/0678-1 · DBLP profile ↗
← Back
15ranked-venue papers
10as first author
13since 2021 · last 2025
0000-0002-2509-1223ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 6 first-author · 6 since 2021Artificial intelligence and machine learning · 7 · 5 first-author · 7 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Enhanced Visual-Semantic Interaction with Tailored Prompts for Pedestrian Attribute Recognition
abstract
Pedestrian attribute recognition (PAR) seeks to predict multiple semantic attributes associated with a specific pedestrian. There are two types of approaches for PAR: unimodal framework and bimodal framework. The former one is to seek a robust visual feature. However, the lack of exploiting semantic feature of linguistic modality is the main concern. The latter one utilizes prompt learning techniques to integrate linguistic data. However, static prompt templates and simple bimodal concatenation cannot to capture the extensive intra-class attribute variability and support active modalities collaboration. In this paper, we propose an Enhanced Visual-Semantic Interaction with Tailored Prompts (EVSITP) framework for PAR. We present an Image-Conditional Dual-Prompt Initialization Module (IDIM) to adaptively generate context-sensitive prompts from visual inputs. Subsequently, a Prompt Enhanced and Regularization Module (PERM) is proposed to strengthen linguistic information from IDIM. We further design a Bimodal Mutual Interaction Module (BMIM) to ensure bidirectional modalities communication. In addition, existing PAR datasets are collected over a short period in limited scenarios, which do not align with real-world scenarios. Therefore, we annotate a long-term person re-identification dataset to create a new PAR dataset, Celeb-PAR. Experiments on several challenging PAR datasets show that our method outperforms state-of-the-art approaches.
Junyi Wu 0001, Yan Huang 0023, Min Gao 0007, Yuzhen Niu, Yuzhong Chen 0001, Qiang Wu 0001
CVPR1
2025 Layered Semi-Second-Order Information Bottleneck and Auxiliary Domain Classification for Person Re-Identification
Anguo Zhang, Junyi Wu 0001, Yueming Gao, Min Gao 0007, Yongduan Song 0001, Sio-Hang Pun
Int. J. Comput. Vis.2
2025 Rethinking attention mechanism for enhanced pedestrian attribute recognition
abstract
Pedestrian Attribute Recognition (PAR) plays a crucial role in various computer vision applications, demanding precise and reliable identification of attributes from pedestrian images. Traditional PAR methods, though effective in leveraging attention mechanisms, often suffer from the lack of direct supervision on attention, leading to potential overfitting and misallocation. This paper introduces a novel and model-agnostic approach, Attention-Aware Regularization (AAR), which rethinks the attention mechanism by integrating causal reasoning to provide direct supervision of attention maps. AAR employs perturbation techniques and a unique optimization objective to assess and refine attention quality, encouraging the model to prioritize attribute-specific regions. Our method demonstrates significant improvement in PAR performance by mitigating the effects of incorrect attention and fostering a more effective attention mechanism. Experiments on standard datasets showcase the superiority of our approach over existing methods, setting a new benchmark for attention-driven PAR models.
Junyi Wu 0001, Yan Huang 0023, Min Gao 0007, Yuzhen Niu, Yuzhong Chen 0001, Qiang Wu 0001
Neurocomputing1
2025 High-order diversity feature learning for pedestrian attribute recognition
abstract
Pedestrian attribute recognition (PAR) involves accurately identifying multiple attributes present in pedestrian images. There are two main approaches for PAR: part-based method and attention-based method. The former relies on existing segmentation or region detection methods to localize body parts and learn corresponding attribute-specific feature from the corresponding regions, where the performance heavily depends on the accuracy of body region localization. The latter adopts the embedded attention modules or transformer attention to exploit detailed feature. However, it can focus on certain body regions but often provide coarse attention, failing to capture fine-grained details, the learned feature may also be interfered with by irrelevant information. Meanwhile, these methods overlook the global contextual information. This work argues for replacing coarse attention with detailed attention and integrating it with global contextual feature from ViT to jointly represent attribute-specific regions. To tackle this issue, we propose a High-order Diversity Feature Learning (HDFL) method for PAR based on ViT. We utilize a polynomial predictor to design an Attribute-specific Detailed Feature Exploration (ADFE) module, which can construct the high-order statistics and gain more fine-grained feature. Our ADFE module is a parameter-friendly method that provides flexibility in deciding its utilization during the inference phase. A Soft-redundancy Perception Loss (SPLoss) is proposed to adaptively measure the redundancy between feature of different orders, which can promote diverse characterization of features. Experiments on several PAR datasets show that our method achieves a new state-of-the-art (SOTA) performance. On the most challenging PA100K dataset, our method outperforms previous SOTA by 1.69% and achieves the highest mA of 84.92%.
Junyi Wu 0001, Yan Huang 0023, Min Gao 0007, Yuzhen Niu, Yuzhong Chen 0001, Qiang Wu 0001
Neural Networks1
2025 Learning Comprehensive Representation via Selective Activation and Dual-Level Orthogonality for Pedestrian Attribute Recognition
abstract
Multi-label Pedestrian Attribute Recognition (PAR) involves identifying a series of semantic attributes in person images. Existing PAR solutions typically rely on CNN as the backbone network to extract pedestrian features. Unfortunately, CNNs process only one adjacent region at a time, resulting in the disappearance of long-range relations between different attribute-specific regions. To address this limitation, we adopt the Vision Transformer (ViT) instead of CNN as the backbone for PAR, aiming to build long-range relations and extract more robust features. However, PAR suffers from an inherent attribute imbalance issue, causing ViT to naturally focus more on attributes that appear frequently in the training set and ignore some pedestrian attributes that appear less. The native features extracted by ViT are not able to tolerate the imbalance attribute distribution issue. To tackle this issue, we propose a novel component and a dual-level loss: the Selective Feature Activation Method (SFAM), the Orthogonal Feature Activation Loss (OFALoss), and Orthogonal Weight Regularization Loss (OWRLoss). SFAM smartly suppresses the more informative attribute-specific features, thus compelling the PAR model to pay greater attention to attribute-specific regions that are often overlooked. The proposed OFALoss enforces an orthogonal constraint on the original feature extracted by ViT and the suppressed features from SFAM, promoting the comprehensiveness of feature representation in each attribute-specific region. Furthermore, OWRLoss is employed for decreasing correlations among entries of the last shared classification layer, which can alleviate the highly correlated of weight vectors caused by non-uniform distribution. This can prevent excessive mutual interference among different attributes during attribute recognition. Our model-agnostic approach is plug-and-play, requiring no additional training parameters in the training process. We conduct experiments on several benchmark PAR datasets, including PETA, PA100K, RAPv1, and RAPv2, demonstrating the effectiveness of our method. Specifically, our method outperforms existing state-of-the-art approaches.
Junyi Wu 0001, Yan Huang 0023, Min Gao 0007, Yuzhen Niu, Yuzhong Chen 0001, Qiang Wu 0001, Jianqiang Zhao
IEEE Trans. Circuits Syst. Video Technol.1
2024 Selective and Orthogonal Feature Activation for Pedestrian Attribute Recognition
abstract
Pedestrian Attribute Recognition (PAR) involves identifying the attributes of individuals in person images. Existing PAR methods typically rely on CNNs as the backbone network to extract pedestrian features. However, CNNs process only one adjacent region at a time, leading to the loss of long-range inter-relations between different attribute-specific regions. To address this limitation, we leverage the Vision Transformer (ViT) instead of CNNs as the backbone for PAR, aiming to model long-range relations and extract more robust features. However, PAR suffers from an inherent attribute imbalance issue, causing ViT to naturally focus more on attributes that appear frequently in the training set and ignore some pedestrian attributes that appear less. The native features extracted by ViT are not able to tolerate the imbalance attribute distribution issue. To tackle this issue, we propose two novel components: the Selective Feature Activation Method (SFAM) and the Orthogonal Feature Activation Loss. SFAM smartly suppresses the more informative attribute-specific features, compelling the PAR model to capture discriminative features from regions that are easily overlooked. The proposed loss enforces an orthogonal constraint on the original feature extracted by ViT and the suppressed features from SFAM, promoting the complementarity of features in space. We conduct experiments on several benchmark PAR datasets, including PETA, PA100K, RAPv1, and RAPv2, demonstrating the effectiveness of our method. Specifically, our method outperforms existing state-of-the-art approaches by GRL, IAA-Caps, ALM, and SSC in terms of mA on the four datasets, respectively.
Junyi Wu 0001, Yan Huang 0023, Min Gao 0007, Yuzhen Niu, Jianqiang Zhao
AAAI1
2024 A Two-Stream Hybrid Convolution-Transformer Network Architecture for Clothing-Change Person Re-Identification
abstract
Long-term (also called Clothing-Change) person re-identification (CC-reID) aims at confirming the identity of pedestrians captured at diverse locations and/or times. Current CC-reID methods heavily rely on ID features learned by the CNN architecture. However, with limited receptive fields, CNN is hard to effectively explore some unique but discriminative ID features (e.g., hair style, tattoo and accessories) from small body regions. Compared with CNN, Transformer has certain merits in exploring more diverse ID-unique features1and retaining more details by the multi-head self-attention design and the removal of down-sampling operation. In this paper, a two-stream hybrid Convolution-Transformer Network (CT-Net) is proposed for CC-reID by combining both CNN and Transformer parallelly in an end-to-end learning scheme. Specifically, CT-Net contains a CNN-based stream (C-Stream) and a Transformer-based stream (T-Stream). Compared with using C-Stream only, T-Stream is used to encourage the C-Stream to explore more detailed ID-unique features when the clothing information is no reliable in CC-reID. Specifically, a Feature Supplement Module (FSM) is proposed to transfer features learned by T-Stream to C-Stream from low-level to high-level for mining more ID-unique feature. In order to further enhance the discriminability2and complementary of ID features learned by our CT-Net, we also introduce a hierarchical supervision with bilinear pooling (HSBP). Experimental results demonstrate that CT-Net performs favorably against the state-of-the-art methods over three CC-reID benchmarks. Meanwhile, CT-Net also demonstrates good generalization ability by achieving comparable performance on traditional person re-ID datasets such as Market-1501 and DukeMTMC-reID.
Junyi Wu 0001, Yan Huang 0023, Min Gao 0007, Jianqiang Zhao, Huiji Zhang, Anguo Zhang
IEEE Trans. Multim.1
2024 Low Latency and Sparse Computing Spiking Neural Networks With Self-Driven Adaptive Threshold Plasticity
abstract
Spiking neural networks (SNNs) have captivated the attention worldwide owing to their compelling advantages in low power consumption, high biological plausibility, and strong robustness. However, the intrinsic latency associated with SNNs during inference poses a significant challenge, impeding their further development and application. This latency is caused by the need for spiking neurons to collect electrical stimuli and generate spikes only when their membrane potential exceeds a firing threshold. Considering the firing threshold plays a crucial role in SNN performance, this article proposes a self-driven adaptive threshold plasticity (SATP) mechanism, wherein neurons autonomously adjust the firing thresholds based on their individual state information using unsupervised learning rules, of which the adjustment is triggered by their own firing events. SATP is based on the principle of maximizing the information contained in the output spike rate distribution of each neuron. This article derives the mathematical expression of SATP and provides extensive experimental results, demonstrating that SATP effectively reduces SNN inference latency, further reduces the computation density while improving computational accuracy, so that SATP facilitates SNN models to be with low latency, sparse computing, and high accuracy.
Anguo Zhang, Jieming Shi 0001, Junyi Wu 0001, Yongcheng Zhou, Wei Yu 0027
IEEE Trans. Neural Networks Learn. Syst.3
2023 Modality Interference Decoupling and Representation Alignment for Caricature-Visual Face Recognition
Junyi Wu 0001, Yan Yan 0002, Xinsheng Du, Huiji Zhang, Jianqiang Zhao
PRCV (1)2
2023 Exponential Information Bottleneck Theory Against Intra-Attribute Variations for Pedestrian Attribute Recognition
abstract
Multi-label pedestrian attribute recognition (PAR) involves assigning multiple attributes to pedestrian images captured by video surveillance cameras. Despite its importance, learning robust attribute-related features for PAR remains a challenge due to the large intra-attribute variations in the image space. These variations, which stem from changes in pedestrian poses, illumination conditions, and background noise, make extracted attribute-related features susceptible to irrelevant information or noise interference. Existing PAR methods rely on body prior extractors or attention mechanisms to locate attribute-correlation regions for extracting robust features. However, these methods may not be robust to intra-attribute variations, which limits their effectiveness. To address this challenge, we propose a novel and flexible PAR framework that leverages the exponential information bottleneck (ExpIB) approach. Our ExpIB-Net uses mutual information compression as the main penalty during the early stage of training, thereby eliminating irrelevant information. As training progresses, the mutual information penalty weakens and the Binary Cross-Entropy Loss (BCELoss) contributes to improving the PAR recognition accuracy. Our method can also be integrated into an attention module to form the AttExpIB-Net, which better handles intra-attribute variations for better performance. Additionally, our model-agnostic ExpIB approach is plug-and-play, requiring no additional computational overhead during inference. Experiments on several challenging PAR datasets show that our method outperforms state-of-the-art approaches.
Junyi Wu 0001, Yan Huang 0023, Min Gao 0007, Jianqiang Zhao, Jieming Shi 0001, Anguo Zhang
IEEE Trans. Inf. Forensics Secur.1
2022 Second-order information bottleneck based spiking neural networks for sEMG recognition
Anguo Zhang, Yuzhen Niu, Yueming Gao, Junyi Wu 0001
Inf. Sci.4
2022 Inter-Attribute awareness for pedestrian attribute recognition
Junyi Wu 0001, Yan Huang 0023, Yating Hong, Jianqiang Zhao, Xinsheng Du
Pattern Recognit.1
2021 Dual-Stream Guided-Learning via a Priori Optimization for Person Re-identification
abstract
The task of person re-identification (re-ID) is to find the same pedestrian across non-overlapping camera views. Generally, the performance of person re-ID can be affected by background clutter. However, existing segmentation algorithms cannot obtain perfect foreground masks to cover the background information clearly. In addition, if the background is completely removed, some discriminative ID-related cues (i.e., backpack or companion) may be lost. In this article, we design a dual-stream network consisting of a Provider Stream (P-Stream) and a Receiver Stream (R-Stream). The R-Stream performs an a priori optimization operation on foreground information. The P-Stream acts as a pusher to guide the R-Stream to concentrate on foreground information and some useful ID-related cues in the background. The proposed dual-stream network can make full use of the a priori optimization and guided-learning strategy to learn encouraging foreground information and some useful ID-related information in the background. Our method achieves Rank-1 accuracy of 95.4% on Market-1501, 89.0% on DukeMTMC-reID, 78.9% on CUHK03 (labeled), and 75.4% on CUHK03 (detected), outperforming state-of-the-art methods.
Junyi Wu 0001, Yan Huang 0023, Qiang Wu 0001, Jianqiang Zhao, Liqin Huang
ACM Trans. Multim. Comput. Commun. Appl.1
2020 Generated Data With Sparse Regularized Multi-Pseudo Label for Person Re-Identification
abstract
Recently, Generative Adversarial Network (GAN) has been adopted to improve person re-identification (person re-ID) performance through data augmentation. However, directly leveraging generated data to train a re-ID model may easily lead to over-fitting issue on these extra data and decrease the generalisability of model to learn true ID-related features from real data. Inspired by the previous approach which assigns multi-pseudo labels on the generated data to reduce the risk of over-fitting, we propose to take sparse regularization into consideration. We attempt to further improve the performance of current re-ID models by using the unlabeled generated data. The proposed Sparse Regularized Multi-Pseudo Label (SRMpL) can effectively prevent the over-fitting issue when some larger weights are assigned to the generated data. Our experiments are carried out on two publicly available person re-ID datasets (e.g., Market-1501 and DukeMTMC-reID). Compared with existing unlabeled generated data re-ID solutions, our approach achieves competitive performance. Two classical re-ID models are used to verify our sparse regularization label on generated data, i.e., an ID-embedding network and a two-stream network.
Liqin Huang, Junyi Wu 0001, Yan Huang 0023, Qiang Wu 0001, Jingsong Xu
IEEE Signal Process. Lett.3
2019 Improving Person Re-Identification Performance Using Body Mask Via Cross-Learning Strategy
abstract
The task of person re-identification (re-id) is to find the same pedestrian across non-overlapping cameras. Normally, the performance of person re-id can be affected by background clutters. However, existing segmentation algorithms are hard to obtain perfect foreground person images. To effectively leverage the body (foreground) cue, and in the meantime pay attention to discriminative information in the background (e.g., companion or vehicle), we propose to use a cross-learning strategy to take both foreground and other discriminative information into account. In addition, since currently existing foreground segmentation result always involves noise, we use Label Smoothing Regularization (LSR) to strengthen the generalization capability during our learning process. In experiments, we pick up two state-of-the-art person re-id methods to verify the effectiveness of our proposed cross-learning strategy. Our experiments are carried out on two publicly available person re-id datasets. Obvious performance improvements can be observed on both datasets.
Junyi Wu 0001, Lingxiang Yao, Yan Huang 0023, Jingsong Xu, Qiang Wu 0001, Liqin Huang
VCIP1