Guoqing Zhang 0002

dblp:27/5832-2 · DBLP profile ↗
← Back
51ranked-venue papers
36as first author
32since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 32 · 24 first-author · 22 since 2021Artificial intelligence and machine learning · 16 · 10 first-author · 8 since 2021Databases, data management, data science and information retrieval · 5 · 3 first-author · 4 since 2021
YearPublicationVenuePosition
2026 Implicit Alignment with Complementary Information for Text-based Person Re-identification
Guoqing Zhang 0002, Yadang Chen, Le Sun 0002, Yulin Cao, Yuhui Zheng
Knowl. Based Syst.1
2026 DC2MNet: Lightweight and Efficient Discrete Cosine Channel Modulation Network for Image Restoration
abstract
Image restoration aims to remove degradation factors (such as blur, snow e.g.) from the damaged image and reconstruct a clean image. Although some methods seek solutions from the frequency domain and are proven to be effective, they are still faced two challenges: (i) Degradation blurs cannot be removed well, and (ii) Inverse transform in frequency domain is computationally expensive. To this end, we propose a lightweight and efficient Discrete Cosine Channel Modulation Network (DC2MNet) for recovering images of multiple degraded conditions from the frequency and spatial perspectives. Specifically, we propose a Discrete Cosine Channel Modulation (DCCM) module to extract the most informative lowest-frequency components of features, and subsequently utilize the channel modulation to reconstruct the global structure of the corresponding feature, avoiding inverse transform in high-dimensional spaces. Furthermore, to effectively remove degradation, we propose a Spatial Mask Modulation (SMM) module to suppress degradation blurs in high-frequency features and emphasize local details that are beneficial to image restoration via pixel-level spatial attention. Finally, we embed the DCCM module and SMM module into the Channel Spatial Modulation Block (CSMB) to form the basic component of DC2MNet, which achieves SOTA performance on various restoration tasks through extensive experiments, including image dehazing, deraining, desnowing and multi-weather restoration. The code and pre-trained models will be open source in this repository.
Guoqing Zhang 0002, Wenxuan Fang 0001, Yupeng Shang, Yuhui Zheng, Weisi Lin
IEEE Trans. Circuits Syst. Video Technol.1
2026 Decoupling Localization and Semantics for Open-Set Object Detection
abstract
Open-set object detection (OSOD) is an important research direction in computer vision, focusing on enhancing a model’s ability to detect unknown categories. Current methods are overly dependent on supervision from known categories, resulting in detection bias that substantially impairs the model’s capacity to recognize unknown classes. In this study, we propose a decoupled localization and semantic OSOD method (DLS-OSOD) that refines supervision granularity to improve unknown category perception. Specifically, to reduce the impact of inaccurate localization on classification, we propose a class-agnostic region proposal network (CA-RPN), which removes the binary classification module in the traditional RPN, allowing the model to focus on region positioning. Furthermore, to mitigate misclassification effects on localization, we design a prototype-based region filtering module (PBF), which constructs a compact prototype space using category semantics during training and pre-filters unknown regions before classification based on region-prototype distance during inference. Additionally, we propose the Unknown Feature Expansion (UFE) and Known Feature Preservation (KFP) modules. UFE enhances supervision for unknown categories by synthesizing unknown category features, improving the model’s ability to detect unknown regions. KFP constrains known-category features through textual anchors, preserving the detection performance of known categories. Experiments on benchmark datasets demonstrate the superior performance of our method.
Guoqing Zhang 0002, Yuhui Zheng
IEEE Trans. Circuits Syst. Video Technol.1
2026 Identity Clue Refinement and Enhancement for Visible-Infrared Person Re-Identification
abstract
Visible-Infrared Person Re-Identification (VI-ReID) is a challenging cross-modal matching task due to significant modality discrepancies. While current methods mainly focus on learning modality-invariant features through unified embedding spaces, they often focus solely on the common discriminative semantics across modalities while disregarding the critical role of modality-specific identity-aware knowledge in discriminative feature learning. To bridge this gap, we propose a novel Identity Clue Refinement and Enhancement (ICRE) network to mine and utilize the implicit discriminative knowledge inherent in modality-specific attributes. Initially, we design a Multi-Perception Feature Refinement (MPFR) module that aggregates shallow features from shared branches, aiming to capture modality-specific attributes that are easily overlooked. Then, we propose a Semantic Distillation Cascade Enhancement (SDCE) module, which distills identity-aware knowledge from the aggregated shallow features and guide the learning of modality-invariant features. Finally, an Identity Clues Guided (ICG) Loss is proposed to alleviate the modality discrepancies within the enhanced features and promote the learning of a diverse representation space. Extensive experiments across multiple public datasets clearly show that our proposed ICRE outperforms existing SOTA methods.
Guoqing Zhang 0002, Zhun Wang, Zhonglin Ye, Yuhui Zheng
IEEE Trans. Multim.1
2025 Single stage weakly supervised semantic segmentation via enhanced patch affinity
Jingjie Jiang, Yuhui Zheng, Guoqing Zhang 0002
Image Vis. Comput.3
2025 Local-enhanced representation for text-based person search
Guoqing Zhang 0002, Yuhao Chen 0002, Yuhui Zheng, Gaven Martin, Ruili Wang 0001
Pattern Recognit.1
2025 Adaptive transformer with Pyramid Fusion for cloth-changing Person Re-Identification
Guoqing Zhang 0002, Jieqiong Zhou, Yuhui Zheng, Gaven J. Martin, Ruili Wang 0001
Pattern Recognit.1
2025 InfinitePerson: Innovating Synthetic Data Creation for Generalization Person Re-Identification
abstract
Recently, large-scale synthetic datasets have effectively alleviated the issue of insufficient person re-identification (Re-ID) datasets. However, synthetic datasets grapple with inherent challenges, including the subpar quality of synthetic pedestrians and single data collection. This paper presents InfinitePerson, a costless pipeline that fully utilizes the infinite generation capability of diffusion models to produce diverse UV texture images and effortlessly constructs high-quality synthetic datasets by simulating a real surveillance network. Specifically, we innovatively propose the utilization of diffusion models to generate high-quality, realistic, and diverse UV texture images to address the limitations of clothing textures. This ensures that our 3D character models have complete clothing texture information and look very similar to real-world pedestrians. Moreover, in response to the challenges in replicating synthetic data collection pipelines, we propose a sub-monitoring network data collection method, which can collect pedestrians data from different viewpoints, backgrounds, and lighting conditions through simple scene layout. Finally, a more scalable and realistic large synthetic dataset called InfinitePerson is created, containing 4,700 identities and 535,636 images. Experimental evidence demonstrates show that models trained on InfinitePerson exhibit superior generalization performance, surpassing those trained on both popular real-world and synthetic person Re-ID datasets. The InfinitePerson project is available athttps://github.com/zhguoqing/InfinitePerson.
Guoqing Zhang 0002, Jin Li 0074, Yuhui Zheng, Ruili Wang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2025 Mask-Aware Hierarchical Aggregation Transformer for Occluded Person Re-Identification
abstract
Occluded person re-identification (Re-ID) is a challenging problem due to the absence of notable discriminative features resulting from incomplete body part images and interference from occluded regions. Recently, some transformer-based methods have demonstrated excellent capabilities in resolving this problem, however these methods are not able to precisely focus on the non-occluded body parts and cannot capture fine-grained local features. To achieve these we propose a Mask-Aware Hierarchical Aggregation TrAnsforMer (MAHATMA) method to enhance occluded person Re-ID. Specifically, we propose a Mask Information Embedding (MIE) module, which directs the model to focus on non-occluded body parts by incorporating the mask semantic information of a human body. Furthermore, to effectively capture fine-grained local features, we propose a Hierarchical Feature Aggregation (HFA) module that mines more exploitable high-quality detail information by aggregating hierarchical image patch representations. To further alleviate the feature loss problem, we propose a Diverse Feature Completion (DFC) module, which is able to complete global features through multi-path feature integration. Extensive experimental evaluations demonstrate that our method exhibits superior performance in dealing with occluded and holistic person datasets.
Guoqing Zhang 0002, Yuhui Zheng, Gaven J. Martin, Ruili Wang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2025 CLIP-Based Multi-Modal Feature Learning for Cloth-Changing Person Re-Identification
abstract
Contrastive Language-Image Pre-training (CLIP) has achieved remarkable results in the field of person re-identification (ReID) due to its excellent cross-modal understanding ability and high scalability. Since the text encoder of CLIP mainly focuses on easy-to-describe attributes such as clothing, and clothing is the main interference factor that reduces the recognition accuracy in cloth-changing person ReID (CC ReID). Consequently, directly applying CLIP to cloth-changing scenario may be difficult to adapt to such dynamic feature changes, thereby affecting the precision of identification. To solve this challenge, we propose a CLIP-based multi-modal feature learning framework (CMFF) for CC ReID. Specifically, we first design a pose-aware identity enhancement module (PIE) to enhance the model's perception of identity-intrinsic information. In this branch, to weaken the interference of clothing information, we apply a ranking loss to minimize the difference between appearance and pose in the feature space. Secondly, we propose a global-local hybrid attention module (GLHA), which fuses head and global features through a cross-attention mechanism, enhancing the global recognition ability of key head information. Finally, considering that existing CLIP-based methods often ignore the potential importance of shallow features, we propose a graph-based multi-layer interactive enhancement module (GMIE), which groups and integrates multi-layer features of the image encoder, aiming to enhance the contextual awareness of multi-scale features. Extensive experiments on multiple popular pedestrian datasets validate the outstanding performance of our proposed CMFF.
Guoqing Zhang 0002, Jieqiong Zhou, Yuhui Zheng, Weisi Lin
IEEE Trans. Image Process.1
2024 Local Feature-Emphasizing Transformer for Cloth-Changing Person Re-identification
Jieqiong Zhou, Guoqing Zhang 0002, Yuhui Zheng, Fuguo Zhang
MMAsia2
2024 Progressive discrepancy elimination for visible-infrared person re-identification
Guoqing Zhang 0002, Zhun Wang, Jieqiong Zhou, Yuhui Zheng
Neurocomputing1
2024 Learning dual attention enhancement feature for visible-infrared person re-identification
Guoqing Zhang 0002, Yinyin Zhang, Yuhao Chen 0002, Yuhui Zheng
J. Vis. Commun. Image Represent.1
2024 SDBAD-Net: A Spatial Dual-Branch Attention Dehazing Network Based on Meta-Former Paradigm
abstract
Image dehazing is an emblematical low-level vision task that aims at restoring haze-free images from haze images. Recently, some methods adopts deep learning techniques to rebuild haze-free images. However, in real-world scenarios, complex degradation of captured images and non-uniform spatial distributions of haze will significantly weaken the generalization ability of these models. Accordingly, we propose a novel Spatial Dual-Branch Attention Dehazing network (SDBAD-Net) based on the Meta-Former paradigm for end-to-end dehazing. Specifically, we firstly design a robust Spatial Dual-Branch Attention (SDBA) module to filter the haze distribution features from different densities, which is suitable for both uniform and non-uniform situations. Secondly, we introduce a Structural Features Supplementary (SFS) module to dynamically fuse the contextual structural features in a nonlinear manner, so as to correct the image distortion caused by the lack of structural details. Finally, the quantitative and qualitative experiments are carried out on two challenging datasets, and the results show that our method outperforms most of state-of-the-art algorithms with fewer parameters and faster speed, especially surpassing FFA-Net with only 50% parameters and 7% computational costs. In addition, we ulteriorly explore its performance on object detection in foggy weather with our model on the challenging Real-world Task-driven Testing Set (RTTS), and the surprising results further prove the robustness and wide-applicability of our method.
Guoqing Zhang 0002, Wenxuan Fang 0001, Yuhui Zheng, Ruili Wang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2023 Multi-level Part-aware Feature Disentangling for Text-based Person Search
abstract
Text-based person search is an important sub-task in cross-modality image retrieval, aiming to capture interested person images by giving textual descriptions. The huge information differences between image and text modalities make this task challenging. Recent methods take local-aligned feature learning strategy into consideration, but lack sufficient mining of more local information. Accordingly, we explore a Multi-level Part-aware Feature Disentangling (MPFD) framework to more fully extract visual and textual representations from multiple angles. Specifically, we introduce a Textual Part-aware Matching (TPM) module into the existing baseline, to disentangle local features for detailed information from both visual and textual part-aware aspects. Besides, in order to fuse multiple local features and improve discrimination of global features, we propose a Multi-level Feature Integration (MFI) module which is capable to perceive the relations between features. We carry out adequate experiments on CUHK-PEDES and ICFG-PEDES datasets to verify our proposed framework, and the results demonstrate that MPFD framework performs favorably against the state-of-the-art methods.
Yuhao Chen 0002, Guoqing Zhang 0002, Yuhui Zheng, Weisi Lin
ICME2
2023 Inter-Intra Camera Identity Learning for Person Re-Identification with Training in Single Camera
abstract
Traditional person re-identification (re-ID) methods generally rely on inter-camera person images to smooth the domain disparities between cameras. However, collecting and annotating a large number of inter-camera identities is extremely difficult and time-consuming, and this makes it hard to deploy person re-ID systems in new locations. To tackle this challenge, this paper studies the single-camera-training (SCT) setting where every person in the training set only appears in one camera. In this work, we design a novel inter-intra camera identity learning (I2CIL) framework to effectively address the SCT person re-ID. Specifically, (i) we design a Dual-Branch Identity Learning (DBIL) network consisting of inter-camera and intra-camera learning branches to learn person ID discriminative information. The former learns camera-irrelevant feature representations by constraining the distance of inter-camera negative sample pairs closer than the distance of intra-camera negative sample pairs. The latter focuses on pulling the distance of intra-camera positive sample pairs closer and pushing the distance of intra-camera negative sample pairs further, partially alleviating weak ID discrimination caused by the lack of inter-camera annotations. (ii) We design a Mixed-Sampling Joint Learning (MSJL) strategy, which is capable to capture inter- and intra-camera samples and independently accomplish the inter- and intra-camera learning tasks at the same time, avoiding the mutual interference between the two tasks. Extensive experiments on two public SCT datasets prove the superiority of the proposed approach.
Guoqing Zhang 0002, Zhiyuan Luo 0003, Weisi Lin, Xuan Jing
ICME1
2023 Complementary networks for person re-identification
Guoqing Zhang 0002, Weisi Lin, Arun Kumar Chandran, Xuan Jing
Inf. Sci.1
2023 Transformer-based global-local feature learning model for occluded person re-identification
Guoqing Zhang 0002, Yuhao Chen 0002, Yuhui Zheng
J. Vis. Commun. Image Represent.1
2023 Camera Contrast Learning for Unsupervised Person Re-Identification
abstract
Unsupervised person re-identification (Re-ID) aims at finding the most informative features from unlabeled person datasets. Some recent approaches adopted camera-aware strategies for model training and have thereby achieved highly promising results. However, these methods simultaneously address intra-ID discrepancies of all cameras and require independent learning under each camera, which increases the complexity of algorithm. To resolve this issue, we present a camera contrast learning framework for unsupervised person Re-ID. Our method first proposes a time-based camera contrastive learning module to facilitate model learning. At each iteration, we follow the time contrast principle to select one camera centroid as proxy of each cluster. By enforcing the samples to converge to positive proxies, the correlation between features and cameras can gradually be reduced. Moreover, we design a 3-dimensional attention module to further reduce intra-ID discrepancies caused by background shifts. By re-weighting each feature map element in a spatial-channel order, our module can exactly find identity-invariant semantic cues from regions of interest in person images, no matter how the background change. Experimental results on several popular datasets prove that our work surpasses existing unsupervised person Re-ID approaches to a remarkable extent. The source codes can be found inhttps://github.com/HongweiZhang97/CCL.
Guoqing Zhang 0002, Weisi Lin, Arun Kumar Chandran, Xuan Jing
IEEE Trans. Circuits Syst. Video Technol.1
2023 Multi-Biometric Unified Network for Cloth-Changing Person Re-Identification
abstract
Person re-identification (re-ID) aims to match the same person across different cameras. However, most existing re-ID methods assume that people wear the same clothes in different views, which limit their performance in identifying target pedestrians who change clothes. Cloth-changing re-ID is a quite challenging problem as clothes occupying a large number of pixels in an image becomes invalid or even misleads information. To tackle this problem, we propose a novel Multi-biometric Unified Network (MBUNet) for learning the robustness of cloth-changing re-ID model by exploiting clothing-independent cues. Specifically, we first introduce a multi-biological feature branch to extract a variety of biological features, such as the head, neck, and shoulders to resist cloth-changing. Then, a differential feature attention module (DFAM) is embedded in this branch, which can extract discriminative fine-grained biological features. Besides, we design a differential recombination on max pooling (DRMP) strategy and simultaneously apply a direction-adaptive graph convolutional layer to mine more robust global and pose features. Finally, we propose a Lightweight Domain Adaptation Module (LDAM) that combines the attention mechanism before and after the waveblock to capture and enhance transferable features across scenarios. To further improve the performance of the model, we also integrate mAP optimization into the objective function of our model for joint training to solve the discrete optimization problem of mAP. Extensive experiments on five cloth-changing re-ID datasets demonstrate the advantages of our proposed MBUNet. The code is available at https://github.com/liyeabc/MBUNet.
Guoqing Zhang 0002, Yuhao Chen 0002, Yuhui Zheng
IEEE Trans. Image Process.1
2022 Multi-Biometric Unified Network for Cloth-Changing Person Re-Identification
abstract
Person re-identification (re-ID) aims at matching the same person across different cameras. Most of the existing meth-ods for re- ID assume that people wear the same clothes on different cameras. However, Cloth-Changing re- ID is a quite challenging problem since people are likely to change clothes as the time span increases. To tackle this problem, a Multi-Biometric Unified Network (MBUNet) is proposed to ex-ploit clothing-unrelated cues. We firstly introduce a multi-biological feature branch that aims at extracting a variety of biological features, such as the head, neck, and shoulders to resist clothing changes. To extract discriminative fine-grained biological features, we embed a differential feature attention module (DFAM) for it. Besides, we adopt differ-ential recombination on max pooling (DRMP) and apply a direction-adaptive graph convolutional layer to extract more robust global features and pose features. Extensive experi-ments on three Cloth-Changing re-ID datasets show the ad-vantages of our proposed MBUNet.
Guoqing Zhang 0002, Yuhao Chen 0002, Yuhui Zheng
ICME1
2022 TIPCB: A simple but effective part-based convolutional baseline for text-based person search
Yuhao Chen 0002, Guoqing Zhang 0002, Yujiang Lu, Yuhui Zheng
Neurocomputing2
2022 Close-set camera style distribution alignment for single camera person re-identification
Guoqing Zhang 0002, Yuhao Chen 0002, Yuhui Zheng
Neurocomputing1
2022 Fine-grained-based multi-feature fusion for occluded person re-identification
Guoqing Zhang 0002, Yuhao Chen 0002, Yuhui Zheng
J. Vis. Commun. Image Represent.1
2022 Illumination Unification for Person Re-Identification
abstract
The performance of person re-identification (re-ID) is easily affected by illumination variations caused by different shooting times, places and cameras. Existing illumination-adaptive methods usually require annotating cross-camera pedestrians on each illumination scale, which is unaffordable for a long-term person retrieval system. The cross-illumination person retrieval problem presents a great challenge for accurate person matching. In this paper, we propose a novel method to tackle this task, which only needs to annotate pedestrians on one illumination scale. Specifically, (i) we propose a novel Illumination Estimation and Restoring framework (IER) to estimate the illumination scale of testing images taken at different illumination conditions and restore them to the illumination scale of training images, such that the disparities between training images with uniform illumination and testing images with varying illuminations are reduced. IER achieves promising results on illumination-adaptive dataset and proving itself a proper baseline for cross-illumination person re-ID. (ii) we propose a Mixed Training strategy using both Original and Reconstructed images (MTOR) to further improve model performance. We generate reconstructed images that are consistent with the original training images in content but more similar to the restored images in style. The reconstructed images are combined with the original training images for supervised training to further reduce the domain gap between original training images and restored testing images. To verify the effectiveness of our method, some simulated illumination-adaptive datasets are constructed with various illumination conditions. Extensive experimental results on the simulated datasets validate the effectiveness of the proposed method. The source code is available athttps://github.com/FadeOrigin/IUReId.
Guoqing Zhang 0002, Zhiyuan Luo 0003, Yuhao Chen 0002, Yuhui Zheng, Weisi Lin
IEEE Trans. Circuits Syst. Video Technol.1
2022 Global Relation-Aware Contrast Learning for Unsupervised Person Re-Identification
abstract
The goal of unsupervised person re-identification (Re-ID) is to use unlabeled person images to learn discriminative features. In recent years, many approaches have adopted clustered pseudo labels to construct proxies for contrastive learning, and have thereby achieved great success. However, existing methods of this kind only utilize local structures within IDs to design their proxies while ignoring the relations between samples of different IDs, which limits the improvement for inter-ID discriminative ability. To resolve this issue, we propose a Global Relation-Aware Contrast Learning (GRACL) method for the task of unsupervised Re-ID. Our method first sets up two proxies for each cluster to capture the inter- and intra-ID relations respectively, which enables us to both effectively increase inter-ID variances and reduce the intra-ID discrepancies. Specifically, the samples that are most different from those in different clusters are selected as inter-ID relation-aware proxies, while those that are least similar to samples from the same clusters are employed as intra-ID relation-aware proxies. With the aid of these proxies, we design both inter- and intra-ID relation-aware contrastive learning modules to facilitate model learning. By pulling each sample close to the positive proxy, we can obtain identity-invariant discriminative features. Experiments on five widely-used Re-ID datasets prove that our GRACL model outperforms current state-of-the-art approaches to a remarkable extent.
Guoqing Zhang 0002, Yuhao Chen 0002, Yuhui Zheng
IEEE Trans. Circuits Syst. Video Technol.2
2021 Reference-Aided Part-Aligned Feature Disentangling for Video Person Re-Identification
abstract
Recently, video-based person re-identification (re-ID) has drawn increasing attention in compute vision community because of its practical application prospects. Due to the inaccurate person detections and pose changes, pedestrian misalignment significantly increases the difficulty of feature extraction and matching. To address this problem, in this paper, we propose a Reference-Aided Part-Aligned (RAPA) framework to disentangle robust features of different parts. Firstly, in order to obtain better references between different videos, a pose-based reference feature learning module is introduced. Secondly, an effective relation-based part feature disentangling module is explored to align frames within each video. By means of using both modules, the informative parts of pedestrian in videos are well aligned and more discriminative feature representation is generated. Comprehensive experiments on three widely-used benchmarks, i.e. iLIDS-VID, PRID-2011 and MARS datasets verify the effectiveness of the proposed framework. Our code will be made publicly available.
Guoqing Zhang 0002, Yuhao Chen 0002, Yuhui Zheng, Yi Wu 0001
ICME1
2021 Low Resolution Information Also Matters: Learning Multi-Resolution Representations for Person Re-Identification
abstract
As a prevailing task in video surveillance and forensics field, person re-identification (re-ID) aims to match person images captured from non-overlapped cameras. In unconstrained scenarios, person images often suffer from the resolution mismatch problem, i.e., Cross-Resolution Person Re-ID. To overcome this problem, most existing methods restore low resolution (LR) images to high resolution (HR) by super-resolution (SR). However, they only focus on the HR feature extraction and ignore the valid information from original LR images. In this work, we explore the influence of resolutions on feature extraction and develop a novel method for cross-resolution person re-ID called Multi-Resolution Representations Joint Learning (MRJL). Our method consists of a Resolution Reconstruction Network (RRN) and a Dual Feature Fusion Network (DFFN). The RRN uses an input image to construct a HR version and a LR version with an encoder and two decoders, while the DFFN adopts a dual-branch structure to generate person representations from multi-resolution images. Comprehensive experiments on five benchmarks verify the superiority of the proposed MRJL over the relevent state-of-the-art methods.
Guoqing Zhang 0002, Yuhao Chen 0002, Weisi Lin, Arun Kumar Chandran, Xuan Jing
IJCAI1
2021 Optimal discriminative feature and dictionary learning for image set classification
Guoqing Zhang 0002, Junchuan Yang, Yuhui Zheng, Zhiyuan Luo 0003
Inf. Sci.1
2021 Hybrid-attention guided network with multiple resolution features for person re-identification
Guoqing Zhang 0002, Junchuan Yang, Yuhui Zheng, Yi Wu 0001, Shengyong Chen
Inf. Sci.1
2021 Cross-view kernel collaborative representation classification for person re-identification
Guoqing Zhang 0002, Tong Jiang, Junchuan Yang, Yuhui Zheng
Multim. Tools Appl.1
2021 Deep High-Resolution Representation Learning for Cross-Resolution Person Re-Identification
abstract
Person re-identification (re-ID) tackles the problem of matching person images with the same identity from different cameras. In practical applications, due to the differences in camera performance and distance between cameras and persons of interest, captured person images usually have various resolutions. This problem, named Cross-Resolution Person Re-identification, presents a great challenge for the accurate person matching. In this paper, we propose a Deep High-Resolution Pseudo-Siamese Framework (PS-HRNet) to solve the above problem. Specifically, we first improve the VDSR by introducing existing channel attention (CA) mechanism and harvest a new module, i.e., VDSR-CA, to restore the resolution of low-resolution images and make full use of the different channel information of feature maps. Then we reform the HRNet by designing a novel representation head, HRNet-ReID, to extract discriminating features. In addition, a pseudo-siamese framework is developed to reduce the difference of feature distributions between low-resolution images and high-resolution images. The experimental results on five cross-resolution person datasets verify the effectiveness of our proposed approach. Compared with the state-of-the-art methods, the proposed PS-HRNet improves the Rank-1 accuracy by 3.4%, 6.2%, 2.5%,1.1% and 4.2% on MLR-Market-1501, MLR-CUHK03, MLR-VIPeR, MLR-DukeMTMC-reID, and CAVIAR datasets, respectively, which demonstrates the superiority of our method in handling the Cross-Resolution Person Re-ID task. Our code is available at https://github.com/zhguoqing.
Guoqing Zhang 0002, Zhicheng Dong 0001, Hao Wang 0101, Yuhui Zheng, Shengyong Chen
IEEE Trans. Image Process.1
2020 Cost-sensitive joint feature and dictionary learning for face recognition
Guoqing Zhang 0002, Fatih Porikli, Huaijiang Sun, Quan-Sen Sun, Guiyu Xia, Yuhui Zheng
Neurocomputing1
2020 Optimal Discriminative Projection for Sparse Representation-Based Classification via Bilevel Optimization
abstract
Recently, sparse representation-based classification (SRC) has been widely studied and has produced state-of-the-art results in various classification tasks. Learning useful and computationally convenient representations from complex redundant and highly variable visual data is crucial for the success of SRC. However, how to find the best feature representation to work with SRC remains an open question. In this paper, we present a novel discriminative projection learning approach with the objective of seeking a projection matrix such that the learned low-dimensional representation can fit SRC well and that it has well discriminant ability. More specifically, we formulate the learning algorithm as a bilevel optimization problem, where the optimization includes an ℓ1-norm minimization problem in its constraints. Through the bilevel optimization model, the relationship between sparse representation and the desired feature projection can be explicitly exploited during the learning process. Therefore, SRC can achieve a better performance in the transformed subspace. The optimization model can be solved by using a stochastic gradient ascent algorithm, and the desired gradient is computed using implicit differentiation. Furthermore, our method can be easily extended to learn a dictionary. The extensive experimental results on a series of benchmark databases show that our method outperforms many state-of-the-art algorithms.
Guoqing Zhang 0002, Huaijiang Sun, Yuhui Zheng, Guiyu Xia, Lei Feng 0003, Quan-Sen Sun
IEEE Trans. Circuits Syst. Video Technol.1
2019 Domain adaptive collaborative representation based classification
Guoqing Zhang 0002, Yuhui Zheng, Guiyu Xia
Multim. Tools Appl.1
2019 Multi-Kernel Coupled Projections for Domain Adaptive Dictionary Learning
abstract
Dictionary learning has produced state-of-the-art results in various classification tasks. However, if the training data have a different distribution than the testing data, the learned sparse representation might not be optimal. Recently, several domain-adaptive dictionary learning (DADL) methods and kernels have been proposed and have achieved impressive performance. However, the performance of these single kernel-based methods heavily depends heavily on the choice of the kernel, and the question of how to combine multiple kernel learning (MKL) with the DADL framework has not been well studied. Motivated by these concerns, in this paper, we propose a multi-kernel domain-adaptive sparse representation-based classification (MK-DASRC) and then use it as a criterion to design a multi-kernel sparse representation-based domain-adaptive discriminative projection method, in which the discriminative features of the data in the two domains are simultaneously learned with the dictionary. The purpose of this method is to maximize the between-class sparse reconstruction residuals of data from both domains, and minimize the within-class sparse reconstruction residuals of data in the low-dimensional subspace. Thus, the resulting representations can satisfactorily fit MK-DASRC and simultaneously display discriminability. Extensive experimental results on a series of benchmark databases show that our method performs better than the state-of-the-art methods.
Yuhui Zheng, Guoqing Zhang 0002, Baihua Xiao, Fu Xiao 0001, Jianwei Zhang 0005
IEEE Trans. Multim.3
2018 Robust image compressive sensing based on m-estimator and nonlocal low-rank regularization
Beijia Chen, Huaijiang Sun, Lei Feng 0003, Guiyu Xia, Guoqing Zhang 0002
Neurocomputing5
2018 Mutually exclusive-KSVD: Learning a discriminative dictionary for hyperspectral image classification
Menglan Xie, Zexuan Ji, Guoqing Zhang 0002, Tao Wang 0020, Quan-Sen Sun
Neurocomputing3
2018 Multiple kernel locality-constrained collaborative representation-based discriminant projection for face recognition
Zhichao Zheng 0002, Huaijiang Sun, Guoqing Zhang 0002
Neurocomputing3
2018 Nonlinear Low-Rank Matrix Completion for Human Motion Recovery
abstract
Human motion capture data has been widely used in many areas, but it involves a complex capture process and the captured data inevitably contains missing data due to the occlusions caused by the actor's body or clothing. Motion recovery, which aims to recover the underlying complete motion sequence from its degraded observation, still remains as a challenging task due to the nonlinear structure and kinematics property embedded in motion data. Low-rank matrix completion based methods have shown promising performance in short-time-missing motion recovery problems. However, low-rank matrix completion, which is designed for linear data, lacks the theoretic guarantee when applied to the recovery of nonlinear motion data. To overcome this drawback, we propose a tailored nonlinear matrix completion model for human motion recovery. Within the model, we first learn a combined low-rank kernel via multiple kernel learning. By exploiting the learned kernel, we embed the motion data into a high dimensional Hilbert space where motion data is of desirable low-rank and we then use the low-rank matrix completion to recover motions. In addition, we add two kinematic constraints to the proposed model to preserve the kinematics property of human motion. Extensive experiment results and comparisons with five other state-of-the-art methods demonstrate the advantage of the proposed method.
Guiyu Xia, Huaijiang Sun, Beijia Chen, Qingshan Liu 0001, Lei Feng 0003, Guoqing Zhang 0002, Renlong Hang
IEEE Trans. Image Process.6
2018 Human Motion Segmentation via Robust Kernel Sparse Subspace Clustering
abstract
Studies on human motion have attracted a lot of attentions. Human motion capture data, which much more precisely records human motion than videos do, has been widely used in many areas. Motion segmentation is an indispensable step for many related applications, but current segmentation methods for motion capture data do not effectively model some important characteristics of motion capture data, such as Riemannian manifold structure and containing non-Gaussian noise. In this paper, we convert the segmentation of motion capture data into a temporal subspace clustering problem. Under the framework of sparse subspace clustering, we propose to use the geodesic exponential kernel to model the Riemannian manifold structure, use correntropy to measure the reconstruction error, use the triangle constraint to guarantee temporal continuity in each cluster and use multi-view reconstruction to extract the relations between different joints. Therefore, exploiting some special characteristics of motion capture data, we propose a new segmentation method, which is robust to non-Gaussian noise, since correntropy is a localized similarity measure. We also develop an efficient optimization algorithm based on block coordinate descent method to solve the proposed model. Our optimization algorithm has a linear complexity while sparse subspace clustering is originally a quadratic problem. Extensive experiment results both on simulated noisy data set and real noisy data set demonstrate the advantage of the proposed method.Studies on human motion have attracted a lot of attentions. Human motion capture data, which much more precisely records human motion than videos do, has been widely used in many areas. Motion segmentation is an indispensable step for many related applications, but current segmentation methods for motion capture data do not effectively model some important characteristics of motion capture data, such as Riemannian manifold structure and containing non-Gaussian noise. In this paper, we convert the segmentation of motion capture data into a temporal subspace clustering problem. Under the framework of sparse subspace clustering, we propose to use the geodesic exponential kernel to model the Riemannian manifold structure, use correntropy to measure the reconstruction error, use the triangle constraint to guarantee temporal continuity in each cluster and use multi-view reconstruction to extract the relations between different joints. Therefore, exploiting some special characteristics of motion capture data, we propose a new segmentation method, which is robust to non-Gaussian noise, since correntropy is a localized similarity measure. We also develop an efficient optimization algorithm based on block coordinate descent method to solve the proposed model. Our optimization algorithm has a linear complexity while sparse subspace clustering is originally a quadratic problem. Extensive experiment results both on simulated noisy data set and real noisy data set demonstrate the advantage of the proposed method.
Guiyu Xia, Huaijiang Sun, Lei Feng 0003, Guoqing Zhang 0002, Yazhou Liu
IEEE Trans. Image Process.4
2017 Dual structural consistency based multi-modal correlation propagation projections for data representation
Hongkun Ji, Quan-Sen Sun, Yun-Hao Yuan 0001, Zexuan Ji, Guoqing Zhang 0002, Lei Feng 0003
Multim. Tools Appl.5
2017 Collaborative probabilistic labels for face recognition from single sample per person
Hongkun Ji, Quan-Sen Sun, Zexuan Ji, Yun-Hao Yuan 0001, Guoqing Zhang 0002
Pattern Recognit.5
2017 Optimal Couple Projections for Domain Adaptive Sparse Representation-Based Classification
abstract
In recent years, sparse representation-based classification (SRC) is one of the most successful methods and has been shown impressive performance in various classification tasks. However, when the training data have a different distribution than the testing data, the learned sparse representation may not be optimal, and the performance of SRC will be degraded significantly. To address this problem, in this paper, we propose an optimal couple projections for domain-adaptive SRC (OCPD-SRC) method, in which the discriminative features of data in the two domains are simultaneously learned with the dictionary that can succinctly represent the training and testing data in the projected space. OCPD-SRC is designed based on the decision rule of SRC, with the objective to learn coupled projection matrices and a common discriminative dictionary such that the between-class sparse reconstruction residuals of data from both domains are maximized, and the within-class sparse reconstruction residuals of data are minimized in the projected low-dimensional space. Thus, the resulting representations can well fit SRC and simultaneously have a better discriminant ability. In addition, our method can be easily extended to multiple domains and can be kernelized to deal with the nonlinear structure of data. The optimal solution for the proposed method can be efficiently obtained following the alternative optimization method. Extensive experimental results on a series of benchmark databases show that our method is better or comparable to many state-of-the-art methods.
Guoqing Zhang 0002, Huaijiang Sun, Fatih Porikli, Yazhou Liu, Quan-Sen Sun
IEEE Trans. Image Process.1
2016 Learning multi-kernel multi-view canonical correlations for image recognition
abstract
canonical correlations (M 2 CCs) framework for subspace learning. In the proposed framework, the input data of each original view are mapped into multiple higher dimensional feature spaces by multiple nonlinear mappings determined by different kernels. This makes M 2 CC can discover multiple kinds of useful information of each original view in the feature spaces. With the framework, we further provide a specific multi-view feature learning method based on direct summation kernel strategy and its regularized version. The experimental results in visual recognition tasks demonstrate the effectiveness and robustness of the proposed method.
Yun-Hao Yuan 0001, Yun Li 0010, Xiaobo Shen 0001, Guoqing Zhang 0002, Quan-Sen Sun
Comput. Vis. Media6
2016 Label propagation based on collaborative representation for face recognition
Guoqing Zhang 0002, Huaijiang Sun, Zexuan Ji, Quan-Sen Sun
Neurocomputing1
2016 Kernel collaborative representation based dictionary learning and discriminative projection
Guoqing Zhang 0002, Huaijiang Sun, Guiyu Xia, Quan-Sen Sun
Neurocomputing1
2016 Human motion recovery jointly utilizing statistical and kinematic information
Guiyu Xia, Huaijiang Sun, Guoqing Zhang 0002, Lei Feng 0003
Inf. Sci.3
2016 Kernel dictionary learning based discriminant analysis
Guoqing Zhang 0002, Huaijiang Sun, Zexuan Ji, Guiyu Xia, Lei Feng 0003, Quan-Sen Sun
J. Vis. Commun. Image Represent.1
2016 Cost-sensitive dictionary learning for face recognition
Guoqing Zhang 0002, Huaijiang Sun, Zexuan Ji, Yun-Hao Yuan 0001, Quan-Sen Sun
Pattern Recognit.1
2016 Multiple Kernel Sparse Representation-Based Orthogonal Discriminative Projection and Its Cost-Sensitive Extension
abstract
Sparse representation-based classification (SRC) has been developed and shown great potential for real-world application. Based on SRC, Yang et al. devised an SRC steered discriminative projection (SRC-DP) method. However, as a linear algorithm, SRC-DP cannot handle the data with highly nonlinear distribution. Kernel sparse representation-based classifier (KSRC) is a non-linear extension of SRC and can remedy the drawback of SRC. KSRC requires the use of a predetermined kernel function and selection of the kernel function and its parameters is difficult. Recently, multiple kernel learning for SRC (MKL-SRC) has been proposed to learn a kernel from a set of base kernels. However, MKL-SRC only considers the within-class reconstruction residual while ignoring the between-class relationship, when learning the kernel weights. In this paper, we propose a novel multiple kernel sparse representation-based classifier, and then we use it as a criterion to design a multiple kernel sparse representation-based orthogonal discriminative projection method. The proposed algorithm aims at learning a projection matrix and a corresponding kernel from the given base kernels such that in the low dimension subspace the between-class reconstruction residual is maximized and the within-class reconstruction residual is minimized. Furthermore, to achieve a minimum overall loss by performing recognition in the learned low-dimensional subspace, we introduce cost information into the dimensionality reduction method. The solutions for the proposed method can be efficiently found based on trace ratio optimization method. Extensive experimental results demonstrate the superiority of the proposed algorithm when compared with the state-of-the-art methods.
Guoqing Zhang 0002, Huaijiang Sun, Guiyu Xia, Quan-Sen Sun
IEEE Trans. Image Process.1