Yan Huang 0023

dblp:75/6434-23 · DBLP profile ↗
← Back
46ranked-venue papers
11as first author
37since 2021 · last 2026
0000-0002-1363-5318ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 25 · 7 first-author · 19 since 2021Artificial intelligence and machine learning · 17 · 6 first-author · 15 since 2021Computer networks · 4 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Security and privacy · 3 · 3 since 2021
YearPublicationVenuePosition
2026 Silhouette generation from identity-level skeleton synthesis for improved gait recognition
Yang Wang 0103, Caifeng Shan, Yan Huang 0023, Liang Wang 0001
Pattern Recognit.4
2025 Enhanced Visual-Semantic Interaction with Tailored Prompts for Pedestrian Attribute Recognition
abstract
Pedestrian attribute recognition (PAR) seeks to predict multiple semantic attributes associated with a specific pedestrian. There are two types of approaches for PAR: unimodal framework and bimodal framework. The former one is to seek a robust visual feature. However, the lack of exploiting semantic feature of linguistic modality is the main concern. The latter one utilizes prompt learning techniques to integrate linguistic data. However, static prompt templates and simple bimodal concatenation cannot to capture the extensive intra-class attribute variability and support active modalities collaboration. In this paper, we propose an Enhanced Visual-Semantic Interaction with Tailored Prompts (EVSITP) framework for PAR. We present an Image-Conditional Dual-Prompt Initialization Module (IDIM) to adaptively generate context-sensitive prompts from visual inputs. Subsequently, a Prompt Enhanced and Regularization Module (PERM) is proposed to strengthen linguistic information from IDIM. We further design a Bimodal Mutual Interaction Module (BMIM) to ensure bidirectional modalities communication. In addition, existing PAR datasets are collected over a short period in limited scenarios, which do not align with real-world scenarios. Therefore, we annotate a long-term person re-identification dataset to create a new PAR dataset, Celeb-PAR. Experiments on several challenging PAR datasets show that our method outperforms state-of-the-art approaches.
Junyi Wu 0001, Yan Huang 0023, Min Gao 0007, Yuzhen Niu, Yuzhong Chen 0001, Qiang Wu 0001
CVPR2
2025 CSFRNet: Integrating Clothing Status Awareness for Long-Term Person Re-identification
Yan Huang 0008, Yan Huang 0023, Zhang Zhang 0001, Qiang Wu 0001, Yi Zhong 0002, Liang Wang 0001
Int. J. Comput. Vis.2
2025 Rethinking attention mechanism for enhanced pedestrian attribute recognition
abstract
Pedestrian Attribute Recognition (PAR) plays a crucial role in various computer vision applications, demanding precise and reliable identification of attributes from pedestrian images. Traditional PAR methods, though effective in leveraging attention mechanisms, often suffer from the lack of direct supervision on attention, leading to potential overfitting and misallocation. This paper introduces a novel and model-agnostic approach, Attention-Aware Regularization (AAR), which rethinks the attention mechanism by integrating causal reasoning to provide direct supervision of attention maps. AAR employs perturbation techniques and a unique optimization objective to assess and refine attention quality, encouraging the model to prioritize attribute-specific regions. Our method demonstrates significant improvement in PAR performance by mitigating the effects of incorrect attention and fostering a more effective attention mechanism. Experiments on standard datasets showcase the superiority of our approach over existing methods, setting a new benchmark for attention-driven PAR models.
Junyi Wu 0001, Yan Huang 0023, Min Gao 0007, Yuzhen Niu, Yuzhong Chen 0001, Qiang Wu 0001
Neurocomputing2
2025 Gait recognition via View-aware Part-wise Attention and Multi-scale Dilated Temporal Extractor
Xu Song, Yang Wang 0103, Yan Huang 0023, Caifeng Shan
Image Vis. Comput.3
2025 Overcoming Data Scarcity in Maritime Radar Target Detection via a Complex-Valued Hybrid Spatiotemporal Network
abstract
Detecting small floating targets on the sea surface has long been a major challenge in radar signal processing. Recently, deep learning (DL) has attracted considerable attention for its potential to improve detection probability. However, its performance heavily relies on the availability of sufficiently labeled datasets, which are often difficult to acquire in complex sea clutter environments. Therefore, this letter introduces the Complex-Valued Hybrid Spatio-Temporal Network (CVHSTNet), a novel maritime radar target detection method designed for low-data scenarios that utilizes time-frequency (TF) representations of radar echoes as inputs. To mitigate the overfitting issue, CVHSTNet is intentionally designed with a shallow architecture, integrating a three-layer complex-valued convolutional neural network (CV-CNN) with a one-layer complex-valued bidirectional long short-term memory network (CV-BiLSTM). Unlike existing real-valued models that overlook phase information, our method operates directly on complex-valued data to capture the complete signal representation. More importantly, this hybrid architecture enables the network to effectively exploit both spatial and temporal characteristics, thereby further enhancing feature representations. Comprehensive experiments on 40 datasets from the IPIX database demonstrate that, with only 50 samples per range cell for training, the proposed method achieves a detection probability exceeding 90% in 37 out of 40 datasets, with a false alarm rate (FAR) of 10−3. To the best of our knowledge, this is the first time a DL-based approach has demonstrated the ability to distinguish between small floating targets and sea clutter under limited labeled radar data conditions.
Ju Wang 0008, Chongyue Wang, Zhaojie Li, Yi Zhong 0002, Yan Huang 0023
IEEE Geosci. Remote. Sens. Lett.6
2025 High-order diversity feature learning for pedestrian attribute recognition
abstract
Pedestrian attribute recognition (PAR) involves accurately identifying multiple attributes present in pedestrian images. There are two main approaches for PAR: part-based method and attention-based method. The former relies on existing segmentation or region detection methods to localize body parts and learn corresponding attribute-specific feature from the corresponding regions, where the performance heavily depends on the accuracy of body region localization. The latter adopts the embedded attention modules or transformer attention to exploit detailed feature. However, it can focus on certain body regions but often provide coarse attention, failing to capture fine-grained details, the learned feature may also be interfered with by irrelevant information. Meanwhile, these methods overlook the global contextual information. This work argues for replacing coarse attention with detailed attention and integrating it with global contextual feature from ViT to jointly represent attribute-specific regions. To tackle this issue, we propose a High-order Diversity Feature Learning (HDFL) method for PAR based on ViT. We utilize a polynomial predictor to design an Attribute-specific Detailed Feature Exploration (ADFE) module, which can construct the high-order statistics and gain more fine-grained feature. Our ADFE module is a parameter-friendly method that provides flexibility in deciding its utilization during the inference phase. A Soft-redundancy Perception Loss (SPLoss) is proposed to adaptively measure the redundancy between feature of different orders, which can promote diverse characterization of features. Experiments on several PAR datasets show that our method achieves a new state-of-the-art (SOTA) performance. On the most challenging PA100K dataset, our method outperforms previous SOTA by 1.69% and achieves the highest mA of 84.92%.
Junyi Wu 0001, Yan Huang 0023, Min Gao 0007, Yuzhen Niu, Yuzhong Chen 0001, Qiang Wu 0001
Neural Networks2
2025 Learning Comprehensive Representation via Selective Activation and Dual-Level Orthogonality for Pedestrian Attribute Recognition
abstract
Multi-label Pedestrian Attribute Recognition (PAR) involves identifying a series of semantic attributes in person images. Existing PAR solutions typically rely on CNN as the backbone network to extract pedestrian features. Unfortunately, CNNs process only one adjacent region at a time, resulting in the disappearance of long-range relations between different attribute-specific regions. To address this limitation, we adopt the Vision Transformer (ViT) instead of CNN as the backbone for PAR, aiming to build long-range relations and extract more robust features. However, PAR suffers from an inherent attribute imbalance issue, causing ViT to naturally focus more on attributes that appear frequently in the training set and ignore some pedestrian attributes that appear less. The native features extracted by ViT are not able to tolerate the imbalance attribute distribution issue. To tackle this issue, we propose a novel component and a dual-level loss: the Selective Feature Activation Method (SFAM), the Orthogonal Feature Activation Loss (OFALoss), and Orthogonal Weight Regularization Loss (OWRLoss). SFAM smartly suppresses the more informative attribute-specific features, thus compelling the PAR model to pay greater attention to attribute-specific regions that are often overlooked. The proposed OFALoss enforces an orthogonal constraint on the original feature extracted by ViT and the suppressed features from SFAM, promoting the comprehensiveness of feature representation in each attribute-specific region. Furthermore, OWRLoss is employed for decreasing correlations among entries of the last shared classification layer, which can alleviate the highly correlated of weight vectors caused by non-uniform distribution. This can prevent excessive mutual interference among different attributes during attribute recognition. Our model-agnostic approach is plug-and-play, requiring no additional training parameters in the training process. We conduct experiments on several benchmark PAR datasets, including PETA, PA100K, RAPv1, and RAPv2, demonstrating the effectiveness of our method. Specifically, our method outperforms existing state-of-the-art approaches.
Junyi Wu 0001, Yan Huang 0023, Min Gao 0007, Yuzhen Niu, Yuzhong Chen 0001, Qiang Wu 0001, Jianqiang Zhao
IEEE Trans. Circuits Syst. Video Technol.2
2025 Visible-Infrared Person Re-Identification With Real-World Label Noise
abstract
In recent years, growing needs for advanced security and traffic management have significantly heightened the prominence of the visible-infrared person re-identification community (VI-ReID), garnering considerable attention. A critical challenge in VI-ReID is the performance degradation attributable to label noise, an issue that becomes even more pronounced in cross-modal scenarios due to an increased likelihood of data confusion. While previous methods have achieved notable successes, they often overlook the complexities of instance-dependent and real-world noise, creating a disconnect from the practical applications of person re-identification. To bridge this gap, our research analyzes the primary sources of label noise in real-world settings, which include a) instantiated identities, b) blurry infrared images, and c) annotators’ errors. In response to these challenges, we develop a Robust Hybrid Loss function (RHL) that enables targeted recognition and retrieval optimization through a more fine-grained division of the noisy dataset. The proposed method categorises data into three sets: clean, obviously noisy, and indistinguishably noisy, with bespoke loss calculations for each category. The identification loss is structured to address the varied nature of these sets specifically. For the retrieval sub-task, we utilize an enhanced triplet loss, adept at handling noisy correspondences. Furthermore, to empirically validate our method, we have re-annotated a real-world dataset, SYSU-Real. Our experiments on SYSU-MM01 and RegDB, conducted under various noise ratios of random and instance-dependent label noise, demonstrate the generalized robustness and effectiveness of our proposed approach.
Ruiheng Zhang 0001, Zhe Cao 0001, Yan Huang 0023, Shuo Yang 0006, Lixin Xu 0001, Min Xu 0001
IEEE Trans. Circuits Syst. Video Technol.3
2025 A Benchmark and Frequency Compression Method for Infrared Few-Shot Object Detection
abstract
Infrared few-shot object detection (IFSOD) aims to detect infrared objects with limited labeled examples. Current infrared datasets, however, suffer from limited diversity in object types and classes, hindering robust evaluation of model generalization on novel classes. To systematically assess dataset quality, we propose metrics for class diversity, instance variability, and object density. By integrating three widely used infrared datasets, we construct the first dataset specifically tailored for IFSOD, increasing instance density to 4.8 (a 1.1 improvement) and expanding the number of classes to 18 (a 5-class increase) compared to the source datasets. Furthermore, frequency analysis of spatial features reveals that sparse annotations introduce spectral bias in the frequency domain. Directly transforming spatial features to the frequency domain, however, mixes background noise with object features, causing spectral leakage and impairing the learning of discriminative features for novel classes. To address these issues, we propose the frequency compression few-shot detection (FC-fsd) method, which incorporates a frequency compression (FC) module. The FC module leverages Discrete Cosine Transform (DCT) within localized windows to reduce spectral leakage and enhance feature clarity. With minimal additional computational overhead, FC-fsd significantly outperforms state-of-the-art methods, achieving nAP50 scores of 28.57 (+13.37) and 35.63 (+2.59) in 1-shot and 2-shot settings, respectively. Our dataset is published athttps://github.com/RuihengZhang/IFSOD-dataset.
Ruiheng Zhang 0001, Biwen Yang, Lixin Xu 0001, Yan Huang 0023, Qi Zhang 0070, Zhizhuo Jiang, Yu Liu 0005
IEEE Trans. Geosci. Remote. Sens.4
2025 Learning Guided Implicit Depth Function With Scale-Aware Feature Fusion
abstract
Recently, the single image super-resolution based on implicit image function is a hot topic, which learns a universal model for arbitrary upsampling scales. By contrast, color-guided depth map super-resolution is less explored based on implicit function learning. The related research faces three questions. First, is it also necessary and applicable to fuse the depth feature and the color feature in the encoder with continuous upsampling scales? Second, is the scale information in the encoder as important as that in the decoder? Third, how to efficiently and effectively model the affinity of location distance and content similarity within cross domains in the decoder? This paper proposes a transformer-based network to answer the above questions, which includes a depth super-resolution branch and a guidance extraction branch. Specifically, in the encoder, the effective implicit cross transformer is designed to fuse the guidance from the color feature with continuous coordinate mapping. In addition, the unrelated guidance is filtered out by correlation evaluation in the high-dimension feature space. Unlike the scale only introduced in the decoder, this paper additionally embeds the scale into the position encoding and the feed-forward network in the encoder to learn the scale-aware feature representation. In the decoder, the high-resolution depth feature is reconstructed by using the internal prior and the external guidance. The internal prior is implemented by implicit self-attention in the depth super-resolution branch, and the external guidance is exploited via implicit cross-attention between both branches. Finally, the above decoded features are complementary to generate the high-resolution depth map. The sufficient experiments on the synthetic and real datasets for in-distribution and out-of-distribution upsampling scales validate the improved performance. The code and the models are public via https://github.com/NaNRan13/GIDF.
Yifan Zuo 0001, Yuming Fang 0001, Jiebin Yan, Wenhui Jiang 0001, Yuxin Peng 0001, Yan Huang 0023
IEEE Trans. Image Process.9
2024 Selective and Orthogonal Feature Activation for Pedestrian Attribute Recognition
abstract
Pedestrian Attribute Recognition (PAR) involves identifying the attributes of individuals in person images. Existing PAR methods typically rely on CNNs as the backbone network to extract pedestrian features. However, CNNs process only one adjacent region at a time, leading to the loss of long-range inter-relations between different attribute-specific regions. To address this limitation, we leverage the Vision Transformer (ViT) instead of CNNs as the backbone for PAR, aiming to model long-range relations and extract more robust features. However, PAR suffers from an inherent attribute imbalance issue, causing ViT to naturally focus more on attributes that appear frequently in the training set and ignore some pedestrian attributes that appear less. The native features extracted by ViT are not able to tolerate the imbalance attribute distribution issue. To tackle this issue, we propose two novel components: the Selective Feature Activation Method (SFAM) and the Orthogonal Feature Activation Loss. SFAM smartly suppresses the more informative attribute-specific features, compelling the PAR model to capture discriminative features from regions that are easily overlooked. The proposed loss enforces an orthogonal constraint on the original feature extracted by ViT and the suppressed features from SFAM, promoting the complementarity of features in space. We conduct experiments on several benchmark PAR datasets, including PETA, PA100K, RAPv1, and RAPv2, demonstrating the effectiveness of our method. Specifically, our method outperforms existing state-of-the-art approaches by GRL, IAA-Caps, ALM, and SSC in terms of mA on the four datasets, respectively.
Junyi Wu 0001, Yan Huang 0023, Min Gao 0007, Yuzhen Niu, Jianqiang Zhao
AAAI2
2024 Attribute-Guided Pedestrian Retrieval: Bridging Person Re-ID with Internal Attribute Variability
abstract
In various domains such as surveillance and smart retail, pedestrian retrieval, centering on person re-identification (Re-ID), plays a pivotal role. Existing Re-ID methodologies often overlook subtle internal attribute variations, which are crucial for accurately identifying individuals with changing appearances. In response, our paper introduces the Attribute-Guided Pedestrian Retrieval (AGPR) task, focusing on integrating specified attributes with query images to refine retrieval results. Although there has been progress in attribute-driven image retrieval, there remains a notable gap in effectively blending robust Re-ID models with intra-class attribute variations. To bridge this gap, we present the Attribute-Guided Transformer-based Pedestrian Retrieval (ATPR) framework. ATPR adeptly merges global ID recognition with local attribute learning, ensuring a co-hesive linkage between the two. Furthermore, to effectively handle the complexity of attribute interconnectivity, ATPR organizes attributes into distinct groups and applies both inter-group correlation and intra-group decorrelation regularizations. Our extensive experiments on a newly estab-lished benchmark using the RAP dataset [32] demonstrate the effectiveness of ATPR within the AGPR paradigm.
Yan Huang 0023, Zhang Zhang 0001, Qiang Wu 0001, Yi Zhong 0002, Liang Wang 0001
CVPR1
2024 Concentrating Estimation Attention: Human Prior Constrained Methods for Robust Classification
Zhe Cao 0001, Shuo Yang 0006, Hongbin Pei, Yan Huang 0023, Yushu Yu, Ruiheng Zhang 0001
PRCV (15)6
2024 CFNet: Conditional filter learning with dynamic noise estimation for real image denoising
Yifan Zuo 0001, Wenhao Yao, Yifeng Zeng, Yuming Fang 0001, Yan Huang 0023, Wenhui Jiang 0001
Knowl. Based Syst.6
2024 Customized meta-dataset for automatic classifier accuracy evaluation
Yan Huang 0023, Zhang Zhang 0001, Yan Huang 0008, Qiang Wu 0001, Yi Zhong 0002, Liang Wang 0001
Pattern Recognit.1
2024 Gait Attribute Recognition: A New Benchmark for Learning Richer Attributes From Human Gait Patterns
abstract
Compared to gait recognition, Gait Attribute Recognition (GAR) is a seldom-investigated problem. However, since gait attribute recognition can provide richer and finer semantic descriptions, it is an indispensable part of building intelligent gait analysis systems. Nonetheless, the types of attributes considered in the existing datasets are very limited. This paper contributes a new benchmark dataset for gait attribute recognition named Multi-Attribute Gait (MA-Gait). Our MA-Gait contains 95 subjects recorded from 12 camera views, resulting in more than 13000 sequences, with 16 attributes labeled, including six attributes that have never been considered in the literature. Moreover, we propose a Multi-Scale Motion Encoder (MSME) to extract robust motion features, and an Attribute-Guided Feature Selection Module (AGFSM) to adaptively capture the most discriminative attribute features from static appearance features and dynamic motion features for different attributes. Our method achieves the best GAR accuracy on the new dataset. Comprehensive experiments show the effectiveness of the proposed method through both quantitative and qualitative evaluations.
Xu Song, Saihui Hou, Yan Huang 0023, Chunshui Cao, Xu Liu 0008, Yongzhen Huang, Caifeng Shan
IEEE Trans. Inf. Forensics Secur.3
2024 Enhancing Person Re-Identification Performance Through In Vivo Learning
abstract
This research investigates the potential of in vivo learning to enhance visual representation learning for image-based person re-identification (re-ID). Compared to traditional self-supervised learning (which require external data), the introduced in vivo learning utilizes supervisory labels generated from pedestrian images to improve re-ID accuracy without relying on external data sources. Three carefully designed in vivo learning tasks, leveraging statistical regularities within images, are proposed without the need for laborious manual annotations. These tasks enable feature extractors to learn more comprehensive and discriminative person representations by jointly modeling various aspects of human biological structure information, contributing to enhanced re-ID performance. Notably, the method seamlessly integrates with existing re-ID frameworks, requiring minimal modifications and no additional data beyond the existing training set. Extensive experiments on diverse datasets, including Market1501, CUHK03-NP, Celeb-reID, Celeb-reid-light, PRCC, and LTCC, demonstrate substantial enhancements in rank-1 precision compared to state-of-the-art methods.
Yan Huang 0008, Yan Huang 0023, Zhang Zhang 0001, Qiang Wu 0001, Yi Zhong 0002, Liang Wang 0001
IEEE Trans. Image Process.2
2024 Meta Clothing Status Calibration for Long-Term Person Re-Identification
abstract
Recent studies have seen significant advancements in the field of long-term person re-identification (LT-reID) through the use of clothing-irrelevant or insensitive features. This work takes the field a step further by addressing a previously unexplored issue, the Clothing Status Distribution Shift (CSDS). CSDS refers to the differing ratios of samples with clothing changes to those without clothing changes between the training and test sets, leading to a decline in LT-reID performance. We establish a connection between the performance of LT-reID and CSDS, and argue that addressing CSDS can improve LT-reID performance. To that end, we propose a novel framework called Meta Clothing Status Calibration (MCSC), which uses meta-learning to optimize the LT-reID model. Specifically, MCSC simulates CSDS between meta-train and meta-test with meta-optimization objectives, optimizing the LT-reID model and making it robust to CSDS. This framework is designed to prevent overfitting and improve the generalization ability of the LT-reID model in the presence of CSDS. Comprehensive evaluations on seven datasets demonstrate that the proposed MCSC framework effectively handles CSDS and improves current state-of-the-art LT-reID methods on several LT-reID benchmarks.
Yan Huang 0023, Qiang Wu 0001, Zhang Zhang 0001, Caifeng Shan, Yan Huang 0008, Yi Zhong 0002, Liang Wang 0001
IEEE Trans. Image Process.1
2024 Illumination Distillation Framework for Nighttime Person Re-Identification and a New Benchmark
abstract
Nighttime person Re-ID (person re-identification in the nighttime) is a very important and challenging task for visual surveillance but it has not been thoroughly investigated. Under the low illumination condition, the performance of person Re-ID methods usually sharply deteriorates. To address the low illumination challenge in nighttime person Re-ID, this paper proposes an Illumination Distillation Framework (IDF), which utilizes illumination enhancement and illumination distillation schemes to promote the learning of Re-ID models. Specifically, IDF consists of a master branch, an illumination enhancement branch, and an illumination distillation module. The master branch is used to extract the features from a nighttime image. The illumination enhancement branch first estimates an enhanced image from the nighttime image using a nonlinear curve mapping method and then extracts the enhanced features. However, nighttime and enhanced features usually contain data noise due to unstable lighting conditions and enhancement failures. To fully exploit the complementary benefits of nighttime and enhanced features while suppressing data noise, we propose an illumination distillation module. In particular, the illumination distillation module fuses the features from two branches through a bottleneck fusion model and then uses the fused features to guide the learning of both branches in a distillation manner. In addition, we build a real-world nighttime person Re-ID dataset, namedNight600, which contains 600 identities captured from different viewpoints and nighttime illumination conditions under complex outdoor environments. Experimental results demonstrate that our IDF can achieve state-of-the-art performance on two nighttime person Re-ID datasets (i.e.,Night600andKnight). We will release our code and dataset athttps://github.com/Alexadlu/IDF.
Andong Lu, Zhang Zhang 0001, Yan Huang 0023, Yifan Zhang 0004, Chenglong Li 0002, Jin Tang 0001, Liang Wang 0001
IEEE Trans. Multim.3
2024 A Two-Stream Hybrid Convolution-Transformer Network Architecture for Clothing-Change Person Re-Identification
abstract
Long-term (also called Clothing-Change) person re-identification (CC-reID) aims at confirming the identity of pedestrians captured at diverse locations and/or times. Current CC-reID methods heavily rely on ID features learned by the CNN architecture. However, with limited receptive fields, CNN is hard to effectively explore some unique but discriminative ID features (e.g., hair style, tattoo and accessories) from small body regions. Compared with CNN, Transformer has certain merits in exploring more diverse ID-unique features1and retaining more details by the multi-head self-attention design and the removal of down-sampling operation. In this paper, a two-stream hybrid Convolution-Transformer Network (CT-Net) is proposed for CC-reID by combining both CNN and Transformer parallelly in an end-to-end learning scheme. Specifically, CT-Net contains a CNN-based stream (C-Stream) and a Transformer-based stream (T-Stream). Compared with using C-Stream only, T-Stream is used to encourage the C-Stream to explore more detailed ID-unique features when the clothing information is no reliable in CC-reID. Specifically, a Feature Supplement Module (FSM) is proposed to transfer features learned by T-Stream to C-Stream from low-level to high-level for mining more ID-unique feature. In order to further enhance the discriminability2and complementary of ID features learned by our CT-Net, we also introduce a hierarchical supervision with bilinear pooling (HSBP). Experimental results demonstrate that CT-Net performs favorably against the state-of-the-art methods over three CC-reID benchmarks. Meanwhile, CT-Net also demonstrates good generalization ability by achieving comparable performance on traditional person re-ID datasets such as Market-1501 and DukeMTMC-reID.
Junyi Wu 0001, Yan Huang 0023, Min Gao 0007, Jianqiang Zhao, Huiji Zhang, Anguo Zhang
IEEE Trans. Multim.2
2023 Geometry-assisted multi-representation view reconstruction network for Light Field image angular super-resolution
Deyang Liu, Zaidong Tong, Yan Huang 0023, Yifan Zuo 0001, Yuming Fang 0001
Knowl. Based Syst.3
2023 Optical flow-assisted multi-level fusion network for Light Field image angular reconstruction
Deyang Liu, Yan Huang 0023, Yuanzhi Wang, Yuming Fang 0001
Signal Process. Image Commun.3
2023 Progressive Sub-Domain Information Mining for Single-Source Generalizable Gait Recognition
abstract
Recent years have witnessed the deployment of fully supervised gait recognition. However, due to domain diversity, gait recognition models designed under the fully supervised condition suffer from poor generalization in unseen domains. How to improve the generalization ability of gait recognition models and enhance their performance on unseen domains is still unexplored in existing gait recognition approaches. This paper investigate the generalizable gait recognition problem and proposes a Progressive Sub-domain Information Mining (PSIM) framework for single-source generalizable gait recognition. During training, PSIM can mine sub-domain information from a single large-scale source domain by differentiating gait features extracted from different people through unsupervised clustering. Then, domain information mitigation loss and domain homogenization loss are introduced to regularize those gait features to be domain insensitive. The above procedures is conducted iteratively until the model converges. Our PSIM framework is model-agnostic, which can directly improve the generalization ability of state-of-the-art gait recognition models without bringing too much complexity in model design. In experiments, our model-agnostic PSIM framework is adopted on several gait recognition models to show its effectiveness in boosting gait recognition performance for the single-source generalizable gait recognition task.
Yang Wang 0103, Yan Huang 0023, Caifeng Shan, Liang Wang 0001
IEEE Trans. Inf. Forensics Secur.2
2023 Exponential Information Bottleneck Theory Against Intra-Attribute Variations for Pedestrian Attribute Recognition
abstract
Multi-label pedestrian attribute recognition (PAR) involves assigning multiple attributes to pedestrian images captured by video surveillance cameras. Despite its importance, learning robust attribute-related features for PAR remains a challenge due to the large intra-attribute variations in the image space. These variations, which stem from changes in pedestrian poses, illumination conditions, and background noise, make extracted attribute-related features susceptible to irrelevant information or noise interference. Existing PAR methods rely on body prior extractors or attention mechanisms to locate attribute-correlation regions for extracting robust features. However, these methods may not be robust to intra-attribute variations, which limits their effectiveness. To address this challenge, we propose a novel and flexible PAR framework that leverages the exponential information bottleneck (ExpIB) approach. Our ExpIB-Net uses mutual information compression as the main penalty during the early stage of training, thereby eliminating irrelevant information. As training progresses, the mutual information penalty weakens and the Binary Cross-Entropy Loss (BCELoss) contributes to improving the PAR recognition accuracy. Our method can also be integrated into an attention module to form the AttExpIB-Net, which better handles intra-attribute variations for better performance. Additionally, our model-agnostic ExpIB approach is plug-and-play, requiring no additional computational overhead during inference. Experiments on several challenging PAR datasets show that our method outperforms state-of-the-art approaches.
Junyi Wu 0001, Yan Huang 0023, Min Gao 0007, Jianqiang Zhao, Jieming Shi 0001, Anguo Zhang
IEEE Trans. Inf. Forensics Secur.2
2023 Multi-Stream Dense View Reconstruction Network for Light Field Image Compression
abstract
Recently, many view synthesis-based methods are proposed for high-efficiency light field (LF) image compression. However, most existing methods fail to recover more texture details on occlusion regions, which reduces the compression efficiency. In this paper, we propose a multi-stream dense view reconstruction network to further improve LF image compression performance. In our method, only sparsely-sampled LF views are transmitted and the rest of the views are reconstructed at the decoder side. During the reconstruction process, we firstly constitute a multi-disparity geometry (MDG) structure based on the decoded sparse LF views, which can reflect abundant disparity characteristics. Subsequently, a multi-stream view reconstruction network (MSVRNet) is put forward to reconstruct a high-quality dense LF image, which consists of a multi-scale feature fusion sub-network, a fusion reconstruction sub-network, and a detail refinement sub-network. The multi-scale feature fusion sub-network can implicitly lean abundant multiscale geometric structure features from the constituted MDG structure. The fusion reconstruction sub-network and the detail refinement sub-network are respectively utilized to fuse the learned multiscale geometric features and restore more texture details, especially for occlusion regions. Moreover, 3D convolutional operations are adopted in the whole reconstruction process, which allow information propagation among the learned multiscale geometric features. Comprehensive experimental results demonstrate the effectiveness of the proposed method. The perceptual quality of reconstructed views and application on depth estimation also demonstrate that the proposed method can keep structural consistency of the reconstructed LF image and recover more texture details.
Deyang Liu, Yan Huang 0023, Yuming Fang 0001, Yifan Zuo 0001, Ping An 0001
IEEE Trans. Multim.2
2022 Dynamic Collaboration Convolution for Robust RGBT Tracking
abstract
Learning powerful representation of individual modality is critical for RGBT tracking. Recent works mainly focus on utilizing multiple convolutions to model feature representations of each modality. However, they usually leverage static convolutions to extract features, which are hard to handle complex input data. To deal with this problem, we propose a dynamic collaboration convolution, named DC-Conv, including a set of static convolutions and a weight-router module, for robust RGBT tracking. In specific, we set four static convolutions to each modality in every layer to model each modality, and design a weight-router module to fuse these static convolutions using learned dynamic weights. Such a dynamic weighting scheme makes the convolutions can be adapted to the variations of input data, and thus greatly improves the tracking performance. In addition, we propose an effective progressive learning algorithm to maximize the role of each convolution to make it capture discriminative representations. We evaluate our method on two public RGBT tracking benchmarks, and the results demonstrate the effectiveness of our tracker against state-of-the-art methods.
Andong Lu, Chenglong Li 0002, Yan Huang 0023, Liang Wang 0001
ICPR4
2022 Inter-Attribute awareness for pedestrian attribute recognition
Junyi Wu 0001, Yan Huang 0023, Yating Hong, Jianqiang Zhao, Xinsheng Du
Pattern Recognit.2
2022 A Climate Adaptation Device-Free Sensing Approach for Target Recognition in Foliage Environments
abstract
Accurate and efficient foliage penetration (FOPEN) target recognition plays a vital role in many mission-critical applications, ranging from civilian to surveillance and military. Recently, device-free sensing (DFS), as an emerging technique, has gained great popularity because it requires no dedicated equipment other than wireless transceivers. Although some DFS-based approaches have been successfully applied in foliage environments, they are vulnerable to climate dynamics and heavily rely on re-labeling large amounts of new data when the weather is altered. To address this issue, a CNN-based weather adaptive target recognition network (WATRNet) is proposed in this paper. Specifically, a lightweight weather conditional normalization (WCN) module is embedded atop each convolutional block to encode inputs under different weather conditions into a shared latent feature space. Under an end-to-end learning manner, the proposed WATRNet first learns knowledge from sufficient labeled data under a certain weather condition to achieve a precise classifier. When applying this model under another weather condition, only the WCN module needs to be retrained using limited new labeled samples to learn weather-invariant features, while the rest convolutional parameters in WATRNet are frozen. Consequently, the domain discrepancy caused by climate variations can be adaptively mitigated with as few relabeled data as possible. Comprehensive evaluations are carried out on a real FOPEN dataset collected under four different weather conditions. Experimental results verify that the presented method can achieve over 90% accuracy, even when it implements from a normal weather condition to another severe weather condition with only small amounts of training samples.
Yi Zhong 0002, Tianqi Bi, Ju Wang 0008, Jie Zeng 0001, Yan Huang 0023, Ting Jiang 0008, Siliang Wu
IEEE Trans. Geosci. Remote. Sens.5
2022 Alleviating Modality Bias Training for Infrared-Visible Person Re-Identification
abstract
The task of infrared-visible person re-identification (IV-reID) is to recognize people across two modalities (i.e., RGB and IR). Existing cutting-edge approaches normally use a pair of images that have the same IDs (i.e., ID-tied cross-modality image pairs) and input them into an ImageNet-trained ResNet50. The ResNet50 backbone model can learn shared features across modalities to tolerate modality discrepancies between RGB and IR. This work will unveil a Modality Bias Training (MBT) problem that is less discussed in IV-reID, which will demonstrate that MBT significantly compromises the performance of IV-reID. Due to MBT, IR information can be overwhelmed by RGB information during training when the ResNet50 model is pretrained based on a large amount of RGB images from ImageNet. Thus, the trained models are more inclined to RGB information. Accordingly, the cross-modality generalization ability of the model is also compromised. To tackle this issue, we present a Dual-level Learning Strategy (DLS) that 1) enforces the focus of the network on ID-exclusive (rather than ID-tied) labels of cross-modality image pairs to mitigate the problem of MBT and 2) introduces third modality data that contain both RGB and IR information to further prevent the information from the IR modality from being overwhelmed during training. Our third modality images are generated by a generative adversarial network. A dynamic ID-exclusive Smooth (dIDeS) label is proposed for the generated third modality data. In experiments, comprehensive experiments are carried out to demonstrate the success of DLS in tackling the MBT issue exposed in IV-reID.
Yan Huang 0023, Qiang Wu 0001, Jingsong Xu, Yi Zhong 0002, Peng Zhang 0057, Zhaoxiang Zhang 0001
IEEE Trans. Multim.1
2021 Clothing Status Awareness for Long-Term Person Re-Identification
abstract
Long-Term person re-identification (LT-reID) exposes extreme challenges because of the longer time gaps between two recording footages where a person is likely to change clothing. There are two types of approaches for LT-reID: biometrics-based approach and data adaptation based approach. The former one is to seek clothing irrelevant biometric features. However, seeking high quality biometric feature is the main concern. The latter one adopts fine-tuning strategy by using data with significant clothing change. However, the performance is compromised when it is applied to cases without clothing change. This work argues that these approaches in fact are not aware of clothing status (i.e., change or no-change) of a pedestrian. Instead, they blindly assume all footages of a pedestrian have different clothes. To tackle this issue, a Regularization via Clothing Status Awareness Network (RCSANet) is proposed to regularize descriptions of a pedestrian by embedding the clothing status awareness. Consequently, the description can be enhanced to maintain the best ID discriminative feature while improving its robustness to real-world LT-reID where both clothing-change case and no-clothing-change case exist. Experiments show that RCSANet performs reasonably well on three LT-reID datasets.
Yan Huang 0023, Qiang Wu 0001, Jingsong Xu, Yi Zhong 0002, Zhaoxiang Zhang 0001
ICCV1
2021 Low data regimes in extreme climates: Foliage penetration personnel detection using a wireless network-based device-free sensing approach
Yi Zhong 0002, Tianqi Bi, Ju Wang 0008, Siliang Wu, Ting Jiang 0008, Yan Huang 0023
Ad Hoc Networks6
2021 Unsupervised Domain Adaptation with Background Shift Mitigating for Person Re-Identification
Yan Huang 0023, Qiang Wu 0001, Jingsong Xu, Yi Zhong 0002, Zhaoxiang Zhang 0001
Int. J. Comput. Vis.1
2021 Multilocation Human Activity Recognition via MIMO-OFDM-Based Wireless Networks: An IoT-Inspired Device-Free Sensing Approach
abstract
Device-free sensing (DFS) is an emerging technology that empowers wireless communication systems with the ability for not only data communication but also smart sensing. By taking advantage of machine-learning technologies, DFS transforms traditional wireless communication networks into intelligent context-aware networks and will open the doors for a myriad of promising 6G-enabled Internet of Things (IoT) applications, ranging from smart home to smart buildings. Although significant progress has been made for human activity recognition at a single location by leveraging this technology, performance at multiple locations has not been fully explored. As far as multilocation activity sensing is concerned, the performance is compromised along with the change of locations and labor-intensive annotation works caused by multilocation. To tackle this issue, an activity decomposition network (ActNet) is presented to decompose the activity information directly from input samples by using the training data from different locations together. Instead of dealing with different locations separately, our ActNet can assemble data from different locations together for training to mitigate the data limitation issue caused by a single location. To achieve this, a multiple-input–multiple-output (MIMO)-orthogonal frequency-division multiplexing (OFDM) technology-based prototype system is utilized to collect data samples at 24 different locations in a cluttered office environment. Especially, for each location, only ten samples of each activity are used for training. Experiments demonstrate that the average classification accuracy is 94.6% across all locations with ensured robustness produced by our method.
Yi Zhong 0002, Ju Wang 0008, Siliang Wu, Ting Jiang 0008, Yan Huang 0023, Qiang Wu 0001
IEEE Internet Things J.5
2021 Learning from EPI-Volume-Stack for Light Field image angular super-resolution
Deyang Liu, Qiang Wu 0001, Yan Huang 0023, Xinpeng Huang, Ping An 0001
Signal Process. Image Commun.3
2021 Learning Spatial-Temporal Representations Over Walking Tracklet for Long-Term Person Re-Identification in the Wild
abstract
Long-term person re-identification (re-ID) aims to build identity correspondence of the Target Subject of Interest (TSI) exposed under surveillance cameras over a long time interval. Compared to the conventional short-term re-ID studied by most existing works, it suffers an additional problem: significant dressing change observed with time lapsing. Unfortunately, this variation in long-term person re-ID case contradicts the assumption of prior short-term re-ID approaches, and thus causes significant difficulties if conventional short-term re-ID methods are applied. To address the problem, this paper proposes to learn hybrid feature representation via a two-stream network named SpTSkM, including a spatial-temporal stream and a skeleton motion stream. The former performs directly on image sequences, which tends to learn identity-related spatial-temporal patterns such as body geometric structure and body movement. The latter operates on normalized 3D skeletons by adapting graph convolutional network, which tends to learn pure motion patterns from skeleton sequences. Both streams extract fine-grained level time-gap stable information that is robust to appearance changes in long-term re-ID and meanwhile maintains sufficient discriminability to differentiate different people. The final matching metric is obtained by mixing information of the two streams in a score-level fusion strategy. In addition, we collect a Cloth-Varying vIDeo re-ID (CVID-reID) dataset particularly for long-term re-ID. It contains video tracklets of celebrities posted on the Internet. These videos are snapshots under extremely different scenarios that include highly dynamic background, diverse camera views and abundant cloth variations on each TSI. These factors cause CVID-reID more complicated and closer to practice. Our experiments demonstrate the difficulty of long-term person re-ID and also validate the effectiveness of the proposed SpTSkM, showing the best performance.
Peng Zhang 0057, Jingsong Xu, Qiang Wu 0001, Yan Huang 0023, Xianye Ben
IEEE Trans. Multim.4
2021 Dual-Stream Guided-Learning via a Priori Optimization for Person Re-identification
abstract
The task of person re-identification (re-ID) is to find the same pedestrian across non-overlapping camera views. Generally, the performance of person re-ID can be affected by background clutter. However, existing segmentation algorithms cannot obtain perfect foreground masks to cover the background information clearly. In addition, if the background is completely removed, some discriminative ID-related cues (i.e., backpack or companion) may be lost. In this article, we design a dual-stream network consisting of a Provider Stream (P-Stream) and a Receiver Stream (R-Stream). The R-Stream performs an a priori optimization operation on foreground information. The P-Stream acts as a pusher to guide the R-Stream to concentrate on foreground information and some useful ID-related cues in the background. The proposed dual-stream network can make full use of the a priori optimization and guided-learning strategy to learn encouraging foreground information and some useful ID-related information in the background. Our method achieves Rank-1 accuracy of 95.4% on Market-1501, 89.0% on DukeMTMC-reID, 78.9% on CUHK03 (labeled), and 75.4% on CUHK03 (detected), outperforming state-of-the-art methods.
Junyi Wu 0001, Yan Huang 0023, Qiang Wu 0001, Jianqiang Zhao, Liqin Huang
ACM Trans. Multim. Comput. Commun. Appl.2
2020 Generated Data With Sparse Regularized Multi-Pseudo Label for Person Re-Identification
abstract
Recently, Generative Adversarial Network (GAN) has been adopted to improve person re-identification (person re-ID) performance through data augmentation. However, directly leveraging generated data to train a re-ID model may easily lead to over-fitting issue on these extra data and decrease the generalisability of model to learn true ID-related features from real data. Inspired by the previous approach which assigns multi-pseudo labels on the generated data to reduce the risk of over-fitting, we propose to take sparse regularization into consideration. We attempt to further improve the performance of current re-ID models by using the unlabeled generated data. The proposed Sparse Regularized Multi-Pseudo Label (SRMpL) can effectively prevent the over-fitting issue when some larger weights are assigned to the generated data. Our experiments are carried out on two publicly available person re-ID datasets (e.g., Market-1501 and DukeMTMC-reID). Compared with existing unlabeled generated data re-ID solutions, our approach achieves competitive performance. Two classical re-ID models are used to verify our sparse regularization label on generated data, i.e., an ID-embedding network and a two-stream network.
Liqin Huang, Junyi Wu 0001, Yan Huang 0023, Qiang Wu 0001, Jingsong Xu
IEEE Signal Process. Lett.4
2020 Beyond Scalar Neuron: Adopting Vector-Neuron Capsules for Long-Term Person Re-Identification
abstract
Current person re-identification (re-ID) works mainly focus on the short-term scenario where a person is less likely to change clothes. However, in the long-term re-ID scenario, a person has a great chance to change clothes. A sophisticated re-ID system should take such changes into account. To facilitate the study of long-term re-ID, this paper introduces a large-scale re-ID dataset called “Celeb-reID” to the community. Unlike previous datasets, the same person can change clothes in the proposed Celeb-reID dataset. Images of Celeb-reID are acquired from the Internet using street snap-shots of celebrities. There is a total of 1,052 IDs with 34,186 images making Celeb-reID being the largest long-term re-ID dataset so far. To tackle the challenge of cloth changes, we propose to use vector-neuron (VN) capsules instead of the traditional scalar neurons (SN) to design our network. Compared with SN, one extra-dimensional information in VN can perceive cloth changes of the same person. We introduce a well-designed ReIDCaps network and integrate capsules to deal with the person re-ID task. Soft Embedding Attention (SEA) and Feature Sparse Representation (FSR) mechanisms are adopted in our network for performance boosting. Experiments are conducted on the proposed long-term re-ID dataset and two common short-term re-ID datasets. Comprehensive analyses are given to demonstrate the challenge exposed in our datasets. Experimental results show that our ReIDCaps can outperform existing state-of-the-art methods by a large margin in the long-term scenario.The new dataset and code will be released to facilitate future researches.
Yan Huang 0023, Jingsong Xu, Qiang Wu 0001, Yi Zhong 0002, Peng Zhang 0057, Zhaoxiang Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2020 Top-Push Constrained Modality-Adaptive Dictionary Learning for Cross-Modality Person Re-Identification
abstract
Person re-identification aims to match person captured by multiple non-overlapping cameras that mainly mean standard RGB cameras. In contemporary surveillance, cameras of different modalities such as infrared cameras and depth cameras are introduced because of their unique advantages in poor illumination scenarios. However, re-identifying the persons across such cameras of different modalities is extremely difficult and, unfortunately, seldom discussed. It is mainly caused by extremely different appearances of the person shown under such different camera modalities. In this paper, we tackle this challenging cross-modality people re-identification through a top-push constrained modality-adaptive dictionary learning. The proposed model asymmetrically projects the heterogeneous features from dissimilar modalities onto a common space. In this way, the modality-specific bias is mitigated. Thus, the heterogeneous data can be simultaneously enforced by a shared dictionary in a canonical space. Moreover, a top-push ranking graph regularization is embedded in the proposed model to improve the discriminability, which efficiently further boosts the matching accuracy. In order to implement the proposed model, an iterative process is developed in this paper to optimize these two processes jointly. Extensive experiments on the benchmark SYSU-MM01 and BIWI RGBD-ID person re-identification datasets show promising results which outperform state-of-the-art methods.
Peng Zhang 0057, Jingsong Xu, Qiang Wu 0001, Yan Huang 0023, Jian Zhang 0002
IEEE Trans. Circuits Syst. Video Technol.4
2019 SBSGAN: Suppression of Inter-Domain Background Shift for Person Re-Identification
abstract
Cross-domain person re-identification (re-ID) is challenging due to the bias between training and testing domains. We observe that if backgrounds in the training and testing datasets are very different, it dramatically introduces difficulties to extract robust pedestrian features, and thus compromises the cross-domain person re-ID performance. In this paper, we formulate such problems as a background shift problem. A Suppression of Background Shift Generative Adversarial Network (SBSGAN) is proposed to generate images with suppressed backgrounds. Unlike simply removing backgrounds using binary masks, SBSGAN allows the generator to decide whether pixels should be preserved or suppressed to reduce segmentation errors caused by noisy foreground masks. Additionally, we take ID-related cues, such as vehicles and companions into consideration. With high-quality generated images, a Densely Associated 2-Stream (DA-2S) network is introduced with Inter Stream Densely Connection (ISDC) modules to strengthen the complementarity of the generated data and ID-related cues. The experiments show that the proposed method achieves competitive performance on three re-ID datasets, i.e., Market-1501, DukeMTMC-reID, and CUHK03, under the cross-domain person re-ID scenario.
Yan Huang 0023, Qiang Wu 0001, Jingsong Xu, Yi Zhong 0002
ICCV1
2019 Celebrities-ReID: A Benchmark for Clothes Variation in Long-Term Person Re-Identification
abstract
This paper considers person re-identification (re-ID) in the case of long-time gap (i.e., long-term re-ID) that concentrates on the challenge of clothes variation of each person. We introduce a new dataset, named Celebrities-reID to handle that challenge. Compared with current datasets, the proposed Celebrities-reID dataset is featured in two aspects. First, it contains 590 persons with 10,842 images, and each person does not wear the same clothing twice, making it the largest clothes variation person re-ID dataset to date. Second, a comprehensive evaluation using state of the arts is carried out to verify the feasibility and new challenge exposed by this dataset. In addition, we propose a benchmark approach to the dataset where a two-step fine-tuning strategy on human body parts is introduced to tackle the challenge of clothes variation. In experiments, we evaluate the feasibility and quality of the proposed Celebrities-reID dataset. The experimental results demonstrate that the proposed benchmark approach is not only able to best tackle clothes variation shown in our dataset but also achieves competitive performance on a widely used person re-ID dataset Market1501, which further proves the reliability of the proposed benchmark approach.
Yan Huang 0023, Qiang Wu 0001, Jingsong Xu, Yi Zhong 0002
IJCNN1
2019 Improving Person Re-Identification Performance Using Body Mask Via Cross-Learning Strategy
abstract
The task of person re-identification (re-id) is to find the same pedestrian across non-overlapping cameras. Normally, the performance of person re-id can be affected by background clutters. However, existing segmentation algorithms are hard to obtain perfect foreground person images. To effectively leverage the body (foreground) cue, and in the meantime pay attention to discriminative information in the background (e.g., companion or vehicle), we propose to use a cross-learning strategy to take both foreground and other discriminative information into account. In addition, since currently existing foreground segmentation result always involves noise, we use Label Smoothing Regularization (LSR) to strengthen the generalization capability during our learning process. In experiments, we pick up two state-of-the-art person re-id methods to verify the effectiveness of our proposed cross-learning strategy. Our experiments are carried out on two publicly available person re-id datasets. Obvious performance improvements can be observed on both datasets.
Junyi Wu 0001, Lingxiang Yao, Yan Huang 0023, Jingsong Xu, Qiang Wu 0001, Liqin Huang
VCIP3
2019 Cost-Effective Foliage Penetration Human Detection Under Severe Weather Conditions Based on Auto-Encoder/Decoder Neural Network
abstract
Military surveillance events and rescue activities are vital missions for the Internet-of-Things. To this end, foliage penetration for human detection plays an important role. However, although the feasibility of that mission has been validated, we observe that it still cannot perform promisingly under severe weather conditions, such as rainy, foggy, and snowy days. Therefore, in this paper, experiments are conducted under severe weather conditions based on a proposed deep learning approach. We present an auto-encoder/decoder (Auto-ED) deep neural network that can learn the deep representation and conduct classification task concurrently. Since the property of cost-effective, the device-free sensing techniques are used to address human detection in our case. As we pursue the signal-based mission, two components are involved in the proposed Auto-ED approach. First, an encoder is utilized that encode signal-based inputs into higher dimensional tensors by fractionally strided convolution operations. Then, a decoder is leveraged with convolution operations to extract deep representations and learn the classifier simultaneously. To verify the effectiveness of the proposed approach, we compare it with several machine learning approaches under different weather conditions. Also, a simulation experiment is conducted by adding additive white Gaussian noise to the original target signals with different signal to noise ratios. Experimental results demonstrate that the proposed approach can best tackle the challenge of human detection under severe weather conditions in the high-clutter foliage environment, which indicates its potential application values in the near future.
Yan Huang 0023, Yi Zhong 0002, Qiang Wu 0001, Eryk Dutkiewicz, Ting Jiang 0008
IEEE Internet Things J.1
2019 Multi-Pseudo Regularized Label for Generated Data in Person Re-Identification
abstract
Sufficient training data normally is required to train deeply learned models. However, due to the expensive manual process for labelling large number of images (i.e., annotation), the amount of available training data (i.e., real data) is always limited. To produce more data for training a deep network, Generative Adversarial Network (GAN) can be used to generate artificial sample data (i.e., generated data). However, the generated data usually does not have annotation labels. To solve this problem, in this paper, we propose a virtual label called Multi-pseudo Regularized Label (MpRL) and assign it to the generated data. With MpRL, the generated data will be used as the supplementary of real training data to train a deep neural network in a semi-supervised learning fashion. To build the corresponding relationship between the real data and generated data, MpRL assigns each generated data a proper virtual label which reflects the likelihood of the affiliation of the generated data to predefined training classes in the real data domain. Unlike the traditional label which usually is a single integral number, the virtual label proposed in this work is a set of weight-based values each individual of which is a number in (0,1] called multi-pseudo label and reflects the degree of relation between each generated data to every pre-defined class of real data. A comprehensive evaluation is carried out by adopting two state-of-the-art convolutional neural networks (CNNs) in our experiments to verify the effectiveness of MpRL. Experiments demonstrate that by assigning MpRL to generated data, we can further improve the person re-ID performance on five re-ID datasets, i.e., Market-1501, DukeMTMC-reID, CUHK03, VIPeR, and CUHK01. The proposed method obtains +6.29%, +6.30%, +5.58%, +5.84%, and +3.48% improvements in rank-1 accuracy over a strong CNN baseline on the five datasets respectively, and outperforms state-of-the-art methods.
Yan Huang 0023, Jingsong Xu, Qiang Wu 0001, Zhedong Zheng, Zhaoxiang Zhang 0001, Jian Zhang 0002
IEEE Trans. Image Process.1
2018 Impact of Seasonal Variations on Foliage Penetration Experiment: A WSN-Based Device-Free Sensing Approach
abstract
Foliage penetration (FOPEN) has been found to be a critical mission for a variety of applications, ranging from surveillance to military. Recently, an emerging technology, namely wireless sensor network (WSN)-based device-free sensing (DFS), has been introduced to the domain of FOPEN. This technology only utilizes radio-frequency signals for target detection and classification; thus, no additional hardware is required, just a wireless transceiver. Although the feasibility of using this technology for human detection indoors has been explored to some extent, it is questionable if the same technology can be transferred to outdoors. As far as FOPEN is concerned, the impact of seasonal variations on detection accuracy can be severe. To address this concern, in this paper, an experiment is conducted in four seasons, and how to ensure reasonable detection accuracy with seasonal variations is intensively investigated. To fully evaluate the potential of using the WSN-based DFS for FOPEN, an impulse-radio ultrawideband technology-based prototype is used to collect data samples in different seasons. Unlike the conventional approach based on a combination of statistical properties of received-signal strength and a support vector machine, this approach adopts two special measures for performance enhancement. One measure is to use a higher order cumulant (HOC) algorithm for feature extraction, so that the impact on detection accuracy due to unwanted clutters can be minimized. The other one is to determine the optimal parameters of the classifier by means of a flower pollination algorithm. Consequently, the adverse effects on detection accuracy due to variations of weather conditions in four seasons can be accommodated. According to the experimental result, it is shown that the average classification accuracy of the presented approach can be improved by at least 20% under all seasons with an ensured robustness.
Yi Zhong 0002, Yang Yang 0034, Xi Zhu 0001, Yan Huang 0023, Eryk Dutkiewicz, Zheng Zhou 0001, Ting Jiang 0008
IEEE Trans. Geosci. Remote. Sens.4