Wei Shi 0009

dblp:44/4066-9 · DBLP profile ↗
← Back
20ranked-venue papers
4as first author
12since 2021 · last 2024
0000-0002-3129-0296ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 16 · 1 first-author · 9 since 2021Artificial intelligence and machine learning · 9 · 3 first-author · 8 since 2021
YearPublicationVenuePosition
2024 Style-Agnostic Representation Learning for Visible-Infrared Person Re-Identification
abstract
One main challenge of visible-infrared person re-identification (VI Re-ID) lies in the large style discrepancy between the heterogeneous data. We present a STyle-Agnostic Representation learning (STAR) framework that bridges the modality gaps at both data and feature levels in a progressive manner. At the data level, we present Cross Modality Blending (CMB), a powerful and parameter-free data augmentation scheme that smoothly synthesizes intermediate modalities by conducting identity-preserving patch exchange and smooth cross-modality blending. At the feature level, we explore the inter-modality feature alignment problem from a new perspective of the style-related feature statistics. Specifically, we design a plug-and-play Adaptive Style Normalization (ASN) module to discard the intrinsic style distractors without losing discriminative content via dual-level adaptive distribution normalization and discriminability compensation. Moreover, considering that an appropriate modality intermediary can convey relevant information on the inter-modality distribution shift, we propose Reciprocal Modality Bridging Learning (RMBL) to better steer the modality bridging process. Two lightweight modality transformation modules are designed in RMBL to model an appropriate intermediate space by manipulating high-order statistics under our shortest distance constraint. Meanwhile, intermediary-guided distribution alignment is reciprocally conducted to align heterogeneous features to the modality intermediary. Experiments on VI Re-ID benchmarks demonstrate the superiority and flexibility of STAR over state-of-the-art methods.
Jianbing Wu, Hong Liu 0008, Wei Shi 0009, Mengyuan Liu 0001, Wenhao Li 0002
IEEE Trans. Multim.3
2023 Boosting Person Re-Identification with Viewpoint Contrastive Learning and Adversarial Training
abstract
Person re-identification (ReID) aims at retrieving a person of interest across multiple cameras. Despite significant progress in person ReID, viewpoint variation remains an obstacle to extracting discriminative features for retrieval. To address this problem, we propose a Viewpoint-Robust Network (VRN) based on contrastive learning and adversarial training to boost person ReID. Specifically, a View-point Confusion (VC) module is proposed to conceal viewpoint information to extract viewpoint-agnostic features. We employ viewpoint contrastive learning to discriminate viewpoints, and then conversely ignore the viewpoint information by adversarial training. Besides, an ID Prototype (IDP) module further enhances the network by introducing a confidence-weighted IDP as a viewpoint-robust ID representation and conducting contrastive metric learning with an IDP triplet loss. Extensive experiments demonstrate the proposed method achieves state-of-the-art performance on widely used datasets Market1501 and MSMT17. Visualization of retrieval results illustrates the effectiveness and robustness of the proposed method.
Xingyue Shi, Hong Liu 0008, Wei Shi 0009, Zihui Zhou, Yidi Li 0001
ICASSP3
2023 Learning Concordant Attention via Target-aware Alignment for Visible-Infrared Person Re-identification
abstract
Owing to the large distribution gap between the heterogeneous data in Visible-Infrared Person Re-identification (VI Re-ID), we point out that existing paradigms often suffer from the inter-modal semantic misalignment issue and thus fail to align and compare local details properly. In this paper, we present Concordant Attention Learning (CAL), a novel framework that learns semantic-aligned representations for VI Re-ID. Specifically, we design the Target-aware Concordant Alignment paradigm, which allows target-aware attention adaptation when aligning heterogeneous samples (i.e., adaptive attention adjustment according to the target image being aligned). This is achieved by exploiting the discriminative clues from the modality counterpart and designing effective modality-agnostic correspondence searching strategies. To ensure semantic concordance during the cross-modal retrieval stage, we further propose MatchDistill, which matches the attention patterns across modalities and learns their underlying semantic correlations by bipartite-graph-based similarity modeling and cross-modal knowledge exchange. Extensive experiments on VI Re-ID benchmark datasets demonstrate the effectiveness and superiority of the proposed CAL.
Jianbing Wu, Hong Liu 0008, Yuxin Su 0004, Wei Shi 0009, Hao Tang 0005
ICCV4
2022 Pose-Guided Feature Disentangling for Occluded Person Re-identification Based on Transformer
abstract
Occluded person re-identification is a challenging task as human body parts could be occluded by some obstacles (e.g. trees, cars, and pedestrians) in certain scenes. Some existing pose-guided methods solve this problem by aligning body parts according to graph matching, but these graph-based methods are not intuitive and complicated. Therefore, we propose a transformer-based Pose-guided Feature Disentangling (PFD) method by utilizing pose information to clearly disentangle semantic components (e.g. human body or joint parts) and selectively match non-occluded parts correspondingly. First, Vision Transformer (ViT) is used to extract the patch features with its strong capability. Second, to preliminarily disentangle the pose information from patch information, the matching and distributing mechanism is leveraged in Pose-guided Feature Aggregation (PFA) module. Third, a set of learnable semantic views are introduced in transformer decoder to implicitly enhance the disentangled body part features. However, those semantic views are not guaranteed to be related to the body without additional supervision. Therefore, Pose-View Matching (PVM) module is proposed to explicitly match visible body parts and automatically separate occlusion features. Fourth, to better prevent the interference of occlusions, we design a Pose-guided Push Loss to emphasize the features of visible body parts. Extensive experiments over five challenging datasets for two tasks (occluded and holistic Re-ID) demonstrate that our proposed PFD is superior promising, which performs favorably against state-of-the-art methods. Code is available at https://github.com/WangTaoAs/PFD_Net
Hong Liu 0008, Pinhao Song, Tianyu Guo 0001, Wei Shi 0009
AAAI5
2022 Unsupervised Domain Adaptation Person Re-Identification by Camera-Aware Style Decoupling and Uncertainty Modeling
abstract
Unsupervised domain adaptation (UDA) person re-identification (re-ID) aims to transfer knowledge learned from labeled source domain to unlabeled target domain and has been successfully applied into a wide range of real-world scenarios. However, existing methods are mainly ineffective at handling domain shift as well as being sensitive to camera styles due to the unannotated target domain. In this paper, we pro-pose a Camera-style Separation and Uncertainty Estimation (CSUE) model to address the problem from two perspectives. To alleviate the negative effect of cross-camera variation, we introduce the Camera-aware Style Decoupling module to im-pose inter-and-intra camera constraints on the feature extracting stage. It can better mine and describe the latent camera invariant features. Moreover, to avoid the inherent defect of clustering, an Uncertainty Modeling module is constructed via estimating the certainty, which helps progressively refine the pseudo labels. Extensive experiments on widely used datasets demonstrate the state-of-the-art performance of our model under the UDA re-ID setting.
Jingwen Guo, Hong Liu 0008, Wei Shi 0009, Hao Tang 0005, Jianbing Wu
ICIP3
2022 Identity-Sensitive Knowledge Propagation for Cloth-Changing Person Re-Identification
abstract
Cloth-changing person re-identification (CC-ReID), which aims to match person identities under clothing changes, is a new rising research topic in recent years. However, typical biometrics-based CC-ReID methods often require cumber-some pose or body part estimators to learn cloth-irrelevant features from human biometric traits, which comes with high computational costs. Besides, the performance is significantly limited due to the resolution degradation of surveillance images. To address the above limitations, we propose an effective Identity-Sensitive Knowledge Propagation framework (DeSKPro) for CC-ReID. Specifically, a Cloth-irrelevant Spatial Attention module is introduced to eliminate the distraction of clothing appearance by acquiring knowledge from the human parsing module. To mitigate the resolution degradation issue and mine identity-sensitive cues from human faces, we propose to restore the missing facial details using prior facial knowledge, which is then propagated to a smaller network. After training, the extra computations for human parsing or face restoration are no longer required. Extensive experiments show that our framework outperforms state-of-the-art methods by a large margin. Our code is available at https://github.com/KimbingNg/DeskPro.
Jianbing Wu, Hong Liu 0008, Wei Shi 0009, Hao Tang 0005, Jingwen Guo
ICIP3
2022 A Cloth-Irrelevant Harmonious Attention Network for Cloth-Changing Person Re-identification
abstract
Cloth-changing person re-identification (CC-ReID) is a challenging task that aims at retrieving the target person across large spatial and temporal spans, with a high probability of changing clothes. To explicitly alleviate the impact of the person changing clothes on re-identification, this paper presents a cloth-irrelevant harmonious attention network (CIHANet) that learns cloth-irrelevant knowledge. Firstly, with the help of human parsing, the color information of human clothing is removed to generate black clothes images. Secondly, the raw person images are used to learn features with more color-based appearance knowledge, while the black clothes images are used to learn features with more cloth-irrelevant knowledge. Then, to fuse the knowledge of two distinct streams, we propose the harmonious attention module, including mutual learning attention and salience guided attention mechanisms. The mutual learning attention mechanism adaptively selects identity-relevant features across feature channels to make two streams interact with each other. The salience guided attention mechanism highlights the cloth-irrelevant areas by transferring the spatial knowledge from the black clothes stream to the raw images stream. Finally, quantitative and qualitative results on three CC-ReID datasets validate the superiority of our method on the CC-ReID task.
Zihui Zhou, Hong Liu 0008, Wei Shi 0009, Hao Tang 0005, Xingyue Shi
ICPR3
2022 IRANet: Identity-relevance aware representation for cloth-changing person re-identification
Wei Shi 0009, Hong Liu 0008, Mengyuan Liu 0001
Image Vis. Comput.1
2022 Image-to-video person re-identification using three-dimensional semantic appearance alignment and cross-modal interactive learning
Wei Shi 0009, Hong Liu 0008, Mengyuan Liu 0001
Pattern Recognit.1
2021 Modality-aware Style Adaptation for RGB-Infrared Person Re-Identification
abstract
RGB-infrared (IR) person re-identification is a challenging task due to the large modality gap between RGB and IR images. Many existing methods bridge the modality gap by style conversion, requiring high-similarity images exchanged by complex CNN structures, like GAN. In this paper, we propose a highly compact modality-aware style adaptation (MSA) framework, which aims to explore more potential relations between RGB and IR modalities by introducing new related modalities. Therefore, the attention is shifted from bridging to filling the modality gap with no requirement on high-quality generated images. To this end, we firstly propose a concise feature-free image generation structure to adapt the original modalities to two new styles that are compatible with both inputs by patch-based pixel redistribution. Secondly, we devise two image style quantification metrics to discriminate styles in image space using luminance and contrast. Thirdly, we design two image-level losses based on the quantified results to guide the style adaptation during an end-to-end four-modality collaborative learning process. Experimental results on two datasets SYSU-MM01 and RegDB show that MSA achieves significant improvements with little extra computation cost and outperforms the state-of-the-art methods.
Ziling Miao, Hong Liu 0008, Wei Shi 0009, Wanlu Xu, Hanrong Ye
IJCAI3
2021 Adversarial Feature Disentanglement for Long-Term Person Re-identification
abstract
Most existing person re-identification methods are effective in short-term scenarios because of their appearance dependencies. However, these methods may fail in long-term scenarios where people might change their clothes. To this end, we propose an adversarial feature disentanglement network (AFD-Net) which contains intra-class reconstruction and inter-class adversary to disentangle the identity-related and identity-unrelated (clothing) features. For intra-class reconstruction, the person images with the same identity are represented and disentangled into identity and clothing features by two separate encoders, and further reconstructed into original images to reduce intra-class feature variations. For inter-class adversary, the disentangled features across different identities are exchanged and recombined to generate adversarial clothes-changing images for training, which makes the identity and clothing features more independent. Especially, to supervise these new generated clothes-changing images, a re-feeding strategy is designed to re-disentangle and reconstruct these new images for image-level self-supervision in the original image space and feature-level soft-supervision in the disentangled feature space. Moreover, we collect a challenging Market-Clothes dataset and a real-world PKU-Market-Reid dataset for evaluation. The results on one large-scale short-term dataset (Market-1501) and five long-term datasets (three public and two we proposed) confirm the superiority of our method against other state-of-the-art methods.
Wanlu Xu, Hong Liu 0008, Wei Shi 0009, Ziling Miao, Zhisheng Lu, Feihu Chen
IJCAI3
2021 PCLoss: Fashion Landmark Estimation with Position Constraint Loss
Meijia Song, Hong Liu 0008, Wei Shi 0009, Xia Li 0005
Pattern Recognit.3
2020 A Fast and Accurate Super-Resolution Network Using Progressive Residual Learning
abstract
Single-image super-resolution (SISR) task has witnessed great strides in the past few years with the development of deep learning. However, most existing studies concentrate on exploiting much deeper super-resolution networks, which are not friendly to the constrained computation resources. In this work, a lightweight network using progressive residual learning for SISR (PRLSR) is proposed to address this issue. Specifically, a progressive residual block (PRB) is designed to progressively downsample deep features for reducing the redundancy and obtaining refined features. Simultaneously, a high-frequency preserving module is proposed to lower the detail loss caused by resolution reduction in PRB. Furthermore, a residual learning-based architecture with learnable weights is utilized to extract multilevel features and adaptively adjust the contribution of residual mapping and identity mapping in residual structure to accelerate convergence. Experimental results on four benchmarks show that our PRLSR achieves superior performance over state-of-the-art methods with a significantly decreased computational cost.
Hong Liu 0008, Zhisheng Lu, Wei Shi 0009, Juanhui Tu
ICASSP3
2020 Position Constraint Loss For Fashion Landmark Estimation
abstract
Fashion landmark estimation aims at locating functional key points of clothes, which has wide potential applications in electronic commerce. However, due to the occlusion and weak outline information, landmark estimation occurs outliers and duplicate detection problems. To alleviate these issues, we propose Position Constraint Loss (PCLoss) to constrain error landmark locations by utilizing the position relationship of landmarks. Specifically, PCLoss adds a regularization term for each landmark to regularize their relative positions, and it can be easily applied to both regression and heatmap based methods without extra computation during inference. Unlike existing approaches that propagate landmark information between feature layers by specific network structures, PCLoss introduces position relations of landmarks in the label space without modifying the network structure. In addition, we leverage the skeleton-like relation of clothing to further strengthen position constraints between landmarks. Extensive experimental results on DeepFashion, FLD and FashionAI demonstrate that our methods can effectively increase the performance of mainstream frameworks by a large margin.
Hong Liu 0008, Meijia Song, Wei Shi 0009, Xia Li 0005
ICASSP3
2020 Robust Audio-Visual Mandarin Speech Recognition Based On Adaptive Decision Fusion And Tone Features
abstract
Audio-visual speech recognition (AVSR) integrates both audio and visual information to perform automatic speech recognition (ASR), which improves the robustness of human-robot interaction systems especially in noise environments. However, few methods and applications have paid attention to AVSR in tonal languages, in which the linguistic feature can play an important role as well as visual information. In this work, we propose a method for AVSR in Mandarin based on adaptive decision fusion as well as making full use of tone features. Firstly, we introduce tone features calculated by Constant Q trasform (CQT) and put them into a CNN-based audio network together with Mel-Frequency Cepstral Coefficient (MFCC) audio features. Then, the visual features are extracted by Discrete Cosine Transform (DCT) from mouth regions in video frames and modeled by an LSTM-based visual network. Finally, an adaptive decision fusion network combines the outputs from both streams to make final predictions. Experimental results on the PKU-AV2 dataset show that the tone features can significantly improve the robustness of Mandarin speech recognition systems, and the adaptability of the proposed method to various noise environments.
Hong Liu 0008, Zhengyan Chen, Wei Shi 0009
ICIP3
2020 Identity-sensitive loss guided and instance feature boosted deep embedding for person search
abstract
Person search aims at detecting and re-identifying pedestrians from whole monitoring images, which is vital for intelligent surveillance. However, this task is still challenging due to the extremely few instances per training identity and inherent fine-grained differences among different identities. To this end, this work proposes an identity-sensitive loss guided and instance feature boosted pipeline to extract deep discriminative feature embedding for person search. First, a prior anchor pre-trained network (PAPN) is designed to obtain proper initial state for the whole deep person search training baseline. Second, a new loss function called instance enhancing loss (IEL) is proposed to learn identity-sensitive features by introducing unlabeled identity information. Specifically, the proposed IEL can selectively utilize unlabeled identities with similar appearances to labeled identities to train the person search network. Third, considering the intra-class compactness of features learned by center loss and contextual inter-class relations, two instance boosting strategies (Boosting) are used to learn more discriminative features. Extensive experiments on two benchmark datasets, namely CUHK-SYSU and PRW, demonstrate the effectiveness of our approach.
Wei Shi 0009, Hong Liu 0008, Mengyuan Liu 0001
Neurocomputing1
2019 Self-Refining Deep Symmetry Enhanced Network for Rain Removal
abstract
Rain removal aims to remove the rain streaks on rain images. The state-of-the-art methods are mostly based on Convolutional Neural Network (CNN). However, as CNN is not equivariant to object rotation, these methods are unsuitable for dealing with the tilted rain streaks. To tackle this problem, we propose Deep Symmetry Enhanced Network (DSEN) that is able to explicitly extract the rotation equivariant features from rain images. In addition, we design a self-refining mechanism to remove the accumulated rain streaks in a coarse-to-fine manner. This mechanism reuses DSEN with a novel information link which passes the gradient flow to the higher stages. Extensive experiments on both synthetic and real-world rain images show that our self-refining DSEN yields the top performance.
Hong Liu 0008, Hanrong Ye, Xia Li 0005, Wei Shi 0009, Mengyuan Liu 0001, Qianru Sun
ICIP4
2018 A Discriminatively Learned Feature Embedding Based on Multi-Loss Fusion For Person Search
abstract
Person search is a challenging task that requires to address pedestrian detection and person re- identification simultaneously. Though significant progress has been made in detection and re-identification respectively, the similar appearances of persons, pedestrian misdetections and false alarms still have adverse effects on person search. To this end, an improved end-to-end person search network with multi -loss is proposed to jointly optimize detection and re-identification. Firstly, a pre-trained network is designed to obtain proper initial state for the whole training network. Then, to enhance the person search model, an improved online instance matching (IOIM) loss is proposed by hardening the distribution of labeled identities and softening the distribution of unlabeled identities. Finally, considering the intra-class compactness of features learned by center loss, the IOIM loss is combined with center loss by the proposed multi-loss fusion strategy, which can learn more discriminative feature embeddings. Experimental results on two challenging datasets CUHK-SYSU and PRW demonstrate our approach significantly outperforms the state-of-the-arts.
Hong Liu 0008, Wei Shi 0009, Weipeng Huang, Qiao Guan
ICASSP2
2018 An End-To-End Siamese Convolutional Neural Network for Loop Closure Detection in Visual Slam System
abstract
Loop closure detection is essential and important in visual simultaneous localization and mapping (SLAM) systems. Most existing methods typically utilize a separate feature extraction part and a similarity metric part. Compared to these methods, an end-to-end network is proposed in this paper to jointly optimize the two parts in a unified framework for further enhancing the interworking between these two parts. First, a two-branch siamese network is designed to learn respective features for each scene of an image pair. Then a hierarchical weighted distance (HWD) layer is proposed to fuse the multi-scale features of each convolutional module and calculate the distance between the image pair. Finally, by using the contrastive loss in the training process, the effective feature representation and similarity metric can be learned simultaneously. Experiments on several open datasets illustrate the superior performance of our approach and demonstrate that the end-to-end network is feasible to conduct the loop closure detection in real time and provides an implementable method for visual SLAM systems.
Hong Liu 0008, Weipeng Huang, Wei Shi 0009
ICASSP4
2018 Instance Enhancing Loss: Deep Identity-Sensitive Feature Embedding for Person Search
abstract
Person search, which is vital for intelligent surveillance, aims at detecting and re-identifying pedestrians from whole monitoring images. However, due to the inaccurate pedestrian detections and extremely few instances per training identity, it remains challenging to learn discriminative representations only by labeled identities for person search. To this end, this paper proposes a novel loss function called instance enhancing loss (IEL) to learn deep identity-sensitive features by introducing unlabeled identity information. Specifically, the proposed IEL can selectively annotate unlabeled identities with similar appearances to labeled identities, and utilize these unlabeled identities in conjunction with labeled identities to train the person search network. The amount of unlabeled identities used as labeled instances can be quantitatively adjusted. Moreover, the proposed IEL is trainable and easy to optimize by back propagation algorithms. Extensive experiments on two benchmark datasets, namely CUHK-SYSU and PRW, show that our method outperforms state-of-the-arts for person search.
Wei Shi 0009, Hong Liu 0008, Fanyang Meng, Weipeng Huang
ICIP1