VLDB 2026 Research / reviewers in the wild / expert
Guoqing Zhang 0002
dblp:27/5832-2
· DBLP profile ↗
51ranked-venue papers
36as first author
32since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 32 · 24 first-author · 22 since 2021Artificial intelligence and machine learning · 16 · 10 first-author · 8 since 2021Databases, data management, data science and information retrieval · 5 · 3 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Implicit Alignment with Complementary Information for Text-based Person Re-identification
Guoqing Zhang 0002, Yadang Chen, Le Sun 0002, Yulin Cao, Yuhui Zheng |
Knowl. Based Syst. | 1 |
| 2026 | DC2MNet: Lightweight and Efficient Discrete Cosine Channel Modulation Network for Image RestorationabstractImage restoration aims to remove degradation factors (such as blur, snow e.g.) from the damaged image and reconstruct a clean image. Although some methods seek solutions from the frequency domain and are proven to be effective, they are still faced two challenges: (i) Degradation blurs cannot be removed well, and (ii) Inverse transform in frequency domain is computationally expensive. To this end, we propose a lightweight and efficient Discrete Cosine Channel Modulation Network (DC2MNet) for recovering images of multiple degraded conditions from the frequency and spatial perspectives. Specifically, we propose a Discrete Cosine Channel Modulation (DCCM) module to extract the most informative lowest-frequency components of features, and subsequently utilize the channel modulation to reconstruct the global structure of the corresponding feature, avoiding inverse transform in high-dimensional spaces. Furthermore, to effectively remove degradation, we propose a Spatial Mask Modulation (SMM) module to suppress degradation blurs in high-frequency features and emphasize local details that are beneficial to image restoration via pixel-level spatial attention. Finally, we embed the DCCM module and SMM module into the Channel Spatial Modulation Block (CSMB) to form the basic component of DC2MNet, which achieves SOTA performance on various restoration tasks through extensive experiments, including image dehazing, deraining, desnowing and multi-weather restoration. The code and pre-trained models will be open source in this repository. Guoqing Zhang 0002, Wenxuan Fang 0001, Yupeng Shang, Yuhui Zheng, Weisi Lin |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2026 | Decoupling Localization and Semantics for Open-Set Object DetectionabstractOpen-set object detection (OSOD) is an important research direction in computer vision, focusing on enhancing a model’s ability to detect unknown categories. Current methods are overly dependent on supervision from known categories, resulting in detection bias that substantially impairs the model’s capacity to recognize unknown classes. In this study, we propose a decoupled localization and semantic OSOD method (DLS-OSOD) that refines supervision granularity to improve unknown category perception. Specifically, to reduce the impact of inaccurate localization on classification, we propose a class-agnostic region proposal network (CA-RPN), which removes the binary classification module in the traditional RPN, allowing the model to focus on region positioning. Furthermore, to mitigate misclassification effects on localization, we design a prototype-based region filtering module (PBF), which constructs a compact prototype space using category semantics during training and pre-filters unknown regions before classification based on region-prototype distance during inference. Additionally, we propose the Unknown Feature Expansion (UFE) and Known Feature Preservation (KFP) modules. UFE enhances supervision for unknown categories by synthesizing unknown category features, improving the model’s ability to detect unknown regions. KFP constrains known-category features through textual anchors, preserving the detection performance of known categories. Experiments on benchmark datasets demonstrate the superior performance of our method. Guoqing Zhang 0002, Yuhui Zheng |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2026 | Identity Clue Refinement and Enhancement for Visible-Infrared Person Re-IdentificationabstractVisible-Infrared Person Re-Identification (VI-ReID) is a challenging cross-modal matching task due to significant modality discrepancies. While current methods mainly focus on learning modality-invariant features through unified embedding spaces, they often focus solely on the common discriminative semantics across modalities while disregarding the critical role of modality-specific identity-aware knowledge in discriminative feature learning. To bridge this gap, we propose a novel Identity Clue Refinement and Enhancement (ICRE) network to mine and utilize the implicit discriminative knowledge inherent in modality-specific attributes. Initially, we design a Multi-Perception Feature Refinement (MPFR) module that aggregates shallow features from shared branches, aiming to capture modality-specific attributes that are easily overlooked. Then, we propose a Semantic Distillation Cascade Enhancement (SDCE) module, which distills identity-aware knowledge from the aggregated shallow features and guide the learning of modality-invariant features. Finally, an Identity Clues Guided (ICG) Loss is proposed to alleviate the modality discrepancies within the enhanced features and promote the learning of a diverse representation space. Extensive experiments across multiple public datasets clearly show that our proposed ICRE outperforms existing SOTA methods. Guoqing Zhang 0002, Zhun Wang, Zhonglin Ye, Yuhui Zheng |
IEEE Trans. Multim. | 1 |
| 2025 | Single stage weakly supervised semantic segmentation via enhanced patch affinity
Jingjie Jiang, Yuhui Zheng, Guoqing Zhang 0002 |
Image Vis. Comput. | 3 |
| 2025 | Local-enhanced representation for text-based person search
Guoqing Zhang 0002, Yuhao Chen 0002, Yuhui Zheng, Gaven Martin, Ruili Wang 0001 |
Pattern Recognit. | 1 |
| 2025 | Adaptive transformer with Pyramid Fusion for cloth-changing Person Re-Identification
Guoqing Zhang 0002, Jieqiong Zhou, Yuhui Zheng, Gaven J. Martin, Ruili Wang 0001 |
Pattern Recognit. | 1 |
| 2025 | InfinitePerson: Innovating Synthetic Data Creation for Generalization Person Re-IdentificationabstractRecently, large-scale synthetic datasets have effectively alleviated the issue of insufficient person re-identification (Re-ID) datasets. However, synthetic datasets grapple with inherent challenges, including the subpar quality of synthetic pedestrians and single data collection. This paper presents InfinitePerson, a costless pipeline that fully utilizes the infinite generation capability of diffusion models to produce diverse UV texture images and effortlessly constructs high-quality synthetic datasets by simulating a real surveillance network. Specifically, we innovatively propose the utilization of diffusion models to generate high-quality, realistic, and diverse UV texture images to address the limitations of clothing textures. This ensures that our 3D character models have complete clothing texture information and look very similar to real-world pedestrians. Moreover, in response to the challenges in replicating synthetic data collection pipelines, we propose a sub-monitoring network data collection method, which can collect pedestrians data from different viewpoints, backgrounds, and lighting conditions through simple scene layout. Finally, a more scalable and realistic large synthetic dataset called InfinitePerson is created, containing 4,700 identities and 535,636 images. Experimental evidence demonstrates show that models trained on InfinitePerson exhibit superior generalization performance, surpassing those trained on both popular real-world and synthetic person Re-ID datasets. The InfinitePerson project is available athttps://github.com/zhguoqing/InfinitePerson. Guoqing Zhang 0002, Jin Li 0074, Yuhui Zheng, Ruili Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Mask-Aware Hierarchical Aggregation Transformer for Occluded Person Re-IdentificationabstractOccluded person re-identification (Re-ID) is a challenging problem due to the absence of notable discriminative features resulting from incomplete body part images and interference from occluded regions. Recently, some transformer-based methods have demonstrated excellent capabilities in resolving this problem, however these methods are not able to precisely focus on the non-occluded body parts and cannot capture fine-grained local features. To achieve these we propose a Mask-Aware Hierarchical Aggregation TrAnsforMer (MAHATMA) method to enhance occluded person Re-ID. Specifically, we propose a Mask Information Embedding (MIE) module, which directs the model to focus on non-occluded body parts by incorporating the mask semantic information of a human body. Furthermore, to effectively capture fine-grained local features, we propose a Hierarchical Feature Aggregation (HFA) module that mines more exploitable high-quality detail information by aggregating hierarchical image patch representations. To further alleviate the feature loss problem, we propose a Diverse Feature Completion (DFC) module, which is able to complete global features through multi-path feature integration. Extensive experimental evaluations demonstrate that our method exhibits superior performance in dealing with occluded and holistic person datasets. Guoqing Zhang 0002, Yuhui Zheng, Gaven J. Martin, Ruili Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | CLIP-Based Multi-Modal Feature Learning for Cloth-Changing Person Re-IdentificationabstractContrastive Language-Image Pre-training (CLIP) has achieved remarkable results in the field of person re-identification (ReID) due to its excellent cross-modal understanding ability and high scalability. Since the text encoder of CLIP mainly focuses on easy-to-describe attributes such as clothing, and clothing is the main interference factor that reduces the recognition accuracy in cloth-changing person ReID (CC ReID). Consequently, directly applying CLIP to cloth-changing scenario may be difficult to adapt to such dynamic feature changes, thereby affecting the precision of identification. To solve this challenge, we propose a CLIP-based multi-modal feature learning framework (CMFF) for CC ReID. Specifically, we first design a pose-aware identity enhancement module (PIE) to enhance the model's perception of identity-intrinsic information. In this branch, to weaken the interference of clothing information, we apply a ranking loss to minimize the difference between appearance and pose in the feature space. Secondly, we propose a global-local hybrid attention module (GLHA), which fuses head and global features through a cross-attention mechanism, enhancing the global recognition ability of key head information. Finally, considering that existing CLIP-based methods often ignore the potential importance of shallow features, we propose a graph-based multi-layer interactive enhancement module (GMIE), which groups and integrates multi-layer features of the image encoder, aiming to enhance the contextual awareness of multi-scale features. Extensive experiments on multiple popular pedestrian datasets validate the outstanding performance of our proposed CMFF. Guoqing Zhang 0002, Jieqiong Zhou, Yuhui Zheng, Weisi Lin |
IEEE Trans. Image Process. | 1 |
| 2024 | Local Feature-Emphasizing Transformer for Cloth-Changing Person Re-identification
Jieqiong Zhou, Guoqing Zhang 0002, Yuhui Zheng, Fuguo Zhang |
MMAsia | 2 |
| 2024 | Progressive discrepancy elimination for visible-infrared person re-identification
Guoqing Zhang 0002, Zhun Wang, Jieqiong Zhou, Yuhui Zheng |
Neurocomputing | 1 |
| 2024 | Learning dual attention enhancement feature for visible-infrared person re-identification
Guoqing Zhang 0002, Yinyin Zhang, Yuhao Chen 0002, Yuhui Zheng |
J. Vis. Commun. Image Represent. | 1 |
| 2024 | SDBAD-Net: A Spatial Dual-Branch Attention Dehazing Network Based on Meta-Former ParadigmabstractImage dehazing is an emblematical low-level vision task that aims at restoring haze-free images from haze images. Recently, some methods adopts deep learning techniques to rebuild haze-free images. However, in real-world scenarios, complex degradation of captured images and non-uniform spatial distributions of haze will significantly weaken the generalization ability of these models. Accordingly, we propose a novel Spatial Dual-Branch Attention Dehazing network (SDBAD-Net) based on the Meta-Former paradigm for end-to-end dehazing. Specifically, we firstly design a robust Spatial Dual-Branch Attention (SDBA) module to filter the haze distribution features from different densities, which is suitable for both uniform and non-uniform situations. Secondly, we introduce a Structural Features Supplementary (SFS) module to dynamically fuse the contextual structural features in a nonlinear manner, so as to correct the image distortion caused by the lack of structural details. Finally, the quantitative and qualitative experiments are carried out on two challenging datasets, and the results show that our method outperforms most of state-of-the-art algorithms with fewer parameters and faster speed, especially surpassing FFA-Net with only 50% parameters and 7% computational costs. In addition, we ulteriorly explore its performance on object detection in foggy weather with our model on the challenging Real-world Task-driven Testing Set (RTTS), and the surprising results further prove the robustness and wide-applicability of our method. Guoqing Zhang 0002, Wenxuan Fang 0001, Yuhui Zheng, Ruili Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | Multi-level Part-aware Feature Disentangling for Text-based Person SearchabstractText-based person search is an important sub-task in cross-modality image retrieval, aiming to capture interested person images by giving textual descriptions. The huge information differences between image and text modalities make this task challenging. Recent methods take local-aligned feature learning strategy into consideration, but lack sufficient mining of more local information. Accordingly, we explore a Multi-level Part-aware Feature Disentangling (MPFD) framework to more fully extract visual and textual representations from multiple angles. Specifically, we introduce a Textual Part-aware Matching (TPM) module into the existing baseline, to disentangle local features for detailed information from both visual and textual part-aware aspects. Besides, in order to fuse multiple local features and improve discrimination of global features, we propose a Multi-level Feature Integration (MFI) module which is capable to perceive the relations between features. We carry out adequate experiments on CUHK-PEDES and ICFG-PEDES datasets to verify our proposed framework, and the results demonstrate that MPFD framework performs favorably against the state-of-the-art methods. Yuhao Chen 0002, Guoqing Zhang 0002, Yuhui Zheng, Weisi Lin |
ICME | 2 |
| 2023 | Inter-Intra Camera Identity Learning for Person Re-Identification with Training in Single CameraabstractTraditional person re-identification (re-ID) methods generally rely on inter-camera person images to smooth the domain disparities between cameras. However, collecting and annotating a large number of inter-camera identities is extremely difficult and time-consuming, and this makes it hard to deploy person re-ID systems in new locations. To tackle this challenge, this paper studies the single-camera-training (SCT) setting where every person in the training set only appears in one camera. In this work, we design a novel inter-intra camera identity learning (I2CIL) framework to effectively address the SCT person re-ID. Specifically, (i) we design a Dual-Branch Identity Learning (DBIL) network consisting of inter-camera and intra-camera learning branches to learn person ID discriminative information. The former learns camera-irrelevant feature representations by constraining the distance of inter-camera negative sample pairs closer than the distance of intra-camera negative sample pairs. The latter focuses on pulling the distance of intra-camera positive sample pairs closer and pushing the distance of intra-camera negative sample pairs further, partially alleviating weak ID discrimination caused by the lack of inter-camera annotations. (ii) We design a Mixed-Sampling Joint Learning (MSJL) strategy, which is capable to capture inter- and intra-camera samples and independently accomplish the inter- and intra-camera learning tasks at the same time, avoiding the mutual interference between the two tasks. Extensive experiments on two public SCT datasets prove the superiority of the proposed approach. Guoqing Zhang 0002, Zhiyuan Luo 0003, Weisi Lin, Xuan Jing |
ICME | 1 |
| 2023 | Complementary networks for person re-identification
Guoqing Zhang 0002, Weisi Lin, Arun Kumar Chandran, Xuan Jing |
Inf. Sci. | 1 |
| 2023 | Transformer-based global-local feature learning model for occluded person re-identification
Guoqing Zhang 0002, Yuhao Chen 0002, Yuhui Zheng |
J. Vis. Commun. Image Represent. | 1 |
| 2023 | Camera Contrast Learning for Unsupervised Person Re-IdentificationabstractUnsupervised person re-identification (Re-ID) aims at finding the most informative features from unlabeled person datasets. Some recent approaches adopted camera-aware strategies for model training and have thereby achieved highly promising results. However, these methods simultaneously address intra-ID discrepancies of all cameras and require independent learning under each camera, which increases the complexity of algorithm. To resolve this issue, we present a camera contrast learning framework for unsupervised person Re-ID. Our method first proposes a time-based camera contrastive learning module to facilitate model learning. At each iteration, we follow the time contrast principle to select one camera centroid as proxy of each cluster. By enforcing the samples to converge to positive proxies, the correlation between features and cameras can gradually be reduced. Moreover, we design a 3-dimensional attention module to further reduce intra-ID discrepancies caused by background shifts. By re-weighting each feature map element in a spatial-channel order, our module can exactly find identity-invariant semantic cues from regions of interest in person images, no matter how the background change. Experimental results on several popular datasets prove that our work surpasses existing unsupervised person Re-ID approaches to a remarkable extent. The source codes can be found inhttps://github.com/HongweiZhang97/CCL. Guoqing Zhang 0002, Weisi Lin, Arun Kumar Chandran, Xuan Jing |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | Multi-Biometric Unified Network for Cloth-Changing Person Re-IdentificationabstractPerson re-identification (re-ID) aims to match the same person across different cameras. However, most existing re-ID methods assume that people wear the same clothes in different views, which limit their performance in identifying target pedestrians who change clothes. Cloth-changing re-ID is a quite challenging problem as clothes occupying a large number of pixels in an image becomes invalid or even misleads information. To tackle this problem, we propose a novel Multi-biometric Unified Network (MBUNet) for learning the robustness of cloth-changing re-ID model by exploiting clothing-independent cues. Specifically, we first introduce a multi-biological feature branch to extract a variety of biological features, such as the head, neck, and shoulders to resist cloth-changing. Then, a differential feature attention module (DFAM) is embedded in this branch, which can extract discriminative fine-grained biological features. Besides, we design a differential recombination on max pooling (DRMP) strategy and simultaneously apply a direction-adaptive graph convolutional layer to mine more robust global and pose features. Finally, we propose a Lightweight Domain Adaptation Module (LDAM) that combines the attention mechanism before and after the waveblock to capture and enhance transferable features across scenarios. To further improve the performance of the model, we also integrate mAP optimization into the objective function of our model for joint training to solve the discrete optimization problem of mAP. Extensive experiments on five cloth-changing re-ID datasets demonstrate the advantages of our proposed MBUNet. The code is available at https://github.com/liyeabc/MBUNet. Guoqing Zhang 0002, Yuhao Chen 0002, Yuhui Zheng |
IEEE Trans. Image Process. | 1 |
| 2022 | Multi-Biometric Unified Network for Cloth-Changing Person Re-IdentificationabstractPerson re-identification (re-ID) aims at matching the same person across different cameras. Most of the existing meth-ods for re- ID assume that people wear the same clothes on different cameras. However, Cloth-Changing re- ID is a quite challenging problem since people are likely to change clothes as the time span increases. To tackle this problem, a Multi-Biometric Unified Network (MBUNet) is proposed to ex-ploit clothing-unrelated cues. We firstly introduce a multi-biological feature branch that aims at extracting a variety of biological features, such as the head, neck, and shoulders to resist clothing changes. To extract discriminative fine-grained biological features, we embed a differential feature attention module (DFAM) for it. Besides, we adopt differ-ential recombination on max pooling (DRMP) and apply a direction-adaptive graph convolutional layer to extract more robust global features and pose features. Extensive experi-ments on three Cloth-Changing re-ID datasets show the ad-vantages of our proposed MBUNet. Guoqing Zhang 0002, Yuhao Chen 0002, Yuhui Zheng |
ICME | 1 |
| 2022 | TIPCB: A simple but effective part-based convolutional baseline for text-based person search
Yuhao Chen 0002, Guoqing Zhang 0002, Yujiang Lu, Yuhui Zheng |
Neurocomputing | 2 |
| 2022 | Close-set camera style distribution alignment for single camera person re-identification
Guoqing Zhang 0002, Yuhao Chen 0002, Yuhui Zheng |
Neurocomputing | 1 |
| 2022 | Fine-grained-based multi-feature fusion for occluded person re-identification
Guoqing Zhang 0002, Yuhao Chen 0002, Yuhui Zheng |
J. Vis. Commun. Image Represent. | 1 |
| 2022 | Illumination Unification for Person Re-IdentificationabstractThe performance of person re-identification (re-ID) is easily affected by illumination variations caused by different shooting times, places and cameras. Existing illumination-adaptive methods usually require annotating cross-camera pedestrians on each illumination scale, which is unaffordable for a long-term person retrieval system. The cross-illumination person retrieval problem presents a great challenge for accurate person matching. In this paper, we propose a novel method to tackle this task, which only needs to annotate pedestrians on one illumination scale. Specifically, (i) we propose a novel Illumination Estimation and Restoring framework (IER) to estimate the illumination scale of testing images taken at different illumination conditions and restore them to the illumination scale of training images, such that the disparities between training images with uniform illumination and testing images with varying illuminations are reduced. IER achieves promising results on illumination-adaptive dataset and proving itself a proper baseline for cross-illumination person re-ID. (ii) we propose a Mixed Training strategy using both Original and Reconstructed images (MTOR) to further improve model performance. We generate reconstructed images that are consistent with the original training images in content but more similar to the restored images in style. The reconstructed images are combined with the original training images for supervised training to further reduce the domain gap between original training images and restored testing images. To verify the effectiveness of our method, some simulated illumination-adaptive datasets are constructed with various illumination conditions. Extensive experimental results on the simulated datasets validate the effectiveness of the proposed method. The source code is available athttps://github.com/FadeOrigin/IUReId. Guoqing Zhang 0002, Zhiyuan Luo 0003, Yuhao Chen 0002, Yuhui Zheng, Weisi Lin |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Global Relation-Aware Contrast Learning for Unsupervised Person Re-IdentificationabstractThe goal of unsupervised person re-identification (Re-ID) is to use unlabeled person images to learn discriminative features. In recent years, many approaches have adopted clustered pseudo labels to construct proxies for contrastive learning, and have thereby achieved great success. However, existing methods of this kind only utilize local structures within IDs to design their proxies while ignoring the relations between samples of different IDs, which limits the improvement for inter-ID discriminative ability. To resolve this issue, we propose a Global Relation-Aware Contrast Learning (GRACL) method for the task of unsupervised Re-ID. Our method first sets up two proxies for each cluster to capture the inter- and intra-ID relations respectively, which enables us to both effectively increase inter-ID variances and reduce the intra-ID discrepancies. Specifically, the samples that are most different from those in different clusters are selected as inter-ID relation-aware proxies, while those that are least similar to samples from the same clusters are employed as intra-ID relation-aware proxies. With the aid of these proxies, we design both inter- and intra-ID relation-aware contrastive learning modules to facilitate model learning. By pulling each sample close to the positive proxy, we can obtain identity-invariant discriminative features. Experiments on five widely-used Re-ID datasets prove that our GRACL model outperforms current state-of-the-art approaches to a remarkable extent. Guoqing Zhang 0002, Yuhao Chen 0002, Yuhui Zheng |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | Reference-Aided Part-Aligned Feature Disentangling for Video Person Re-IdentificationabstractRecently, video-based person re-identification (re-ID) has drawn increasing attention in compute vision community because of its practical application prospects. Due to the inaccurate person detections and pose changes, pedestrian misalignment significantly increases the difficulty of feature extraction and matching. To address this problem, in this paper, we propose a Reference-Aided Part-Aligned (RAPA) framework to disentangle robust features of different parts. Firstly, in order to obtain better references between different videos, a pose-based reference feature learning module is introduced. Secondly, an effective relation-based part feature disentangling module is explored to align frames within each video. By means of using both modules, the informative parts of pedestrian in videos are well aligned and more discriminative feature representation is generated. Comprehensive experiments on three widely-used benchmarks, i.e. iLIDS-VID, PRID-2011 and MARS datasets verify the effectiveness of the proposed framework. Our code will be made publicly available. Guoqing Zhang 0002, Yuhao Chen 0002, Yuhui Zheng, Yi Wu 0001 |
ICME | 1 |
| 2021 | Low Resolution Information Also Matters: Learning Multi-Resolution Representations for Person Re-IdentificationabstractAs a prevailing task in video surveillance and forensics field, person re-identification (re-ID) aims to match person images captured from non-overlapped cameras. In unconstrained scenarios, person images often suffer from the resolution mismatch problem, i.e., Cross-Resolution Person Re-ID. To overcome this problem, most existing methods restore low resolution (LR) images to high resolution (HR) by super-resolution (SR). However, they only focus on the HR feature extraction and ignore the valid information from original LR images. In this work, we explore the influence of resolutions on feature extraction and develop a novel method for cross-resolution person re-ID called Multi-Resolution Representations Joint Learning (MRJL). Our method consists of a Resolution Reconstruction Network (RRN) and a Dual Feature Fusion Network (DFFN). The RRN uses an input image to construct a HR version and a LR version with an encoder and two decoders, while the DFFN adopts a dual-branch structure to generate person representations from multi-resolution images. Comprehensive experiments on five benchmarks verify the superiority of the proposed MRJL over the relevent state-of-the-art methods. Guoqing Zhang 0002, Yuhao Chen 0002, Weisi Lin, Arun Kumar Chandran, Xuan Jing |
IJCAI | 1 |
| 2021 | Optimal discriminative feature and dictionary learning for image set classification
Guoqing Zhang 0002, Junchuan Yang, Yuhui Zheng, Zhiyuan Luo 0003 |
Inf. Sci. | 1 |
| 2021 | Hybrid-attention guided network with multiple resolution features for person re-identification
Guoqing Zhang 0002, Junchuan Yang, Yuhui Zheng, Yi Wu 0001, Shengyong Chen |
Inf. Sci. | 1 |
| 2021 | Cross-view kernel collaborative representation classification for person re-identification
Guoqing Zhang 0002, Tong Jiang, Junchuan Yang, Yuhui Zheng |
Multim. Tools Appl. | 1 |
| 2021 | Deep High-Resolution Representation Learning for Cross-Resolution Person Re-IdentificationabstractPerson re-identification (re-ID) tackles the problem of matching person images with the same identity from different cameras. In practical applications, due to the differences in camera performance and distance between cameras and persons of interest, captured person images usually have various resolutions. This problem, named Cross-Resolution Person Re-identification, presents a great challenge for the accurate person matching. In this paper, we propose a Deep High-Resolution Pseudo-Siamese Framework (PS-HRNet) to solve the above problem. Specifically, we first improve the VDSR by introducing existing channel attention (CA) mechanism and harvest a new module, i.e., VDSR-CA, to restore the resolution of low-resolution images and make full use of the different channel information of feature maps. Then we reform the HRNet by designing a novel representation head, HRNet-ReID, to extract discriminating features. In addition, a pseudo-siamese framework is developed to reduce the difference of feature distributions between low-resolution images and high-resolution images. The experimental results on five cross-resolution person datasets verify the effectiveness of our proposed approach. Compared with the state-of-the-art methods, the proposed PS-HRNet improves the Rank-1 accuracy by 3.4%, 6.2%, 2.5%,1.1% and 4.2% on MLR-Market-1501, MLR-CUHK03, MLR-VIPeR, MLR-DukeMTMC-reID, and CAVIAR datasets, respectively, which demonstrates the superiority of our method in handling the Cross-Resolution Person Re-ID task. Our code is available at https://github.com/zhguoqing. Guoqing Zhang 0002, Zhicheng Dong 0001, Hao Wang 0101, Yuhui Zheng, Shengyong Chen |
IEEE Trans. Image Process. | 1 |
| 2020 | Cost-sensitive joint feature and dictionary learning for face recognition
Guoqing Zhang 0002, Fatih Porikli, Huaijiang Sun, Quan-Sen Sun, Guiyu Xia, Yuhui Zheng |
Neurocomputing | 1 |
| 2020 | Optimal Discriminative Projection for Sparse Representation-Based Classification via Bilevel OptimizationabstractRecently, sparse representation-based classification (SRC) has been widely studied and has produced state-of-the-art results in various classification tasks. Learning useful and computationally convenient representations from complex redundant and highly variable visual data is crucial for the success of SRC. However, how to find the best feature representation to work with SRC remains an open question. In this paper, we present a novel discriminative projection learning approach with the objective of seeking a projection matrix such that the learned low-dimensional representation can fit SRC well and that it has well discriminant ability. More specifically, we formulate the learning algorithm as a bilevel optimization problem, where the optimization includes an ℓ1-norm minimization problem in its constraints. Through the bilevel optimization model, the relationship between sparse representation and the desired feature projection can be explicitly exploited during the learning process. Therefore, SRC can achieve a better performance in the transformed subspace. The optimization model can be solved by using a stochastic gradient ascent algorithm, and the desired gradient is computed using implicit differentiation. Furthermore, our method can be easily extended to learn a dictionary. The extensive experimental results on a series of benchmark databases show that our method outperforms many state-of-the-art algorithms. Guoqing Zhang 0002, Huaijiang Sun, Yuhui Zheng, Guiyu Xia, Lei Feng 0003, Quan-Sen Sun |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2019 | Domain adaptive collaborative representation based classification
Guoqing Zhang 0002, Yuhui Zheng, Guiyu Xia |
Multim. Tools Appl. | 1 |
| 2019 | Multi-Kernel Coupled Projections for Domain Adaptive Dictionary LearningabstractDictionary learning has produced state-of-the-art results in various classification tasks. However, if the training data have a different distribution than the testing data, the learned sparse representation might not be optimal. Recently, several domain-adaptive dictionary learning (DADL) methods and kernels have been proposed and have achieved impressive performance. However, the performance of these single kernel-based methods heavily depends heavily on the choice of the kernel, and the question of how to combine multiple kernel learning (MKL) with the DADL framework has not been well studied. Motivated by these concerns, in this paper, we propose a multi-kernel domain-adaptive sparse representation-based classification (MK-DASRC) and then use it as a criterion to design a multi-kernel sparse representation-based domain-adaptive discriminative projection method, in which the discriminative features of the data in the two domains are simultaneously learned with the dictionary. The purpose of this method is to maximize the between-class sparse reconstruction residuals of data from both domains, and minimize the within-class sparse reconstruction residuals of data in the low-dimensional subspace. Thus, the resulting representations can satisfactorily fit MK-DASRC and simultaneously display discriminability. Extensive experimental results on a series of benchmark databases show that our method performs better than the state-of-the-art methods. Yuhui Zheng, Guoqing Zhang 0002, Baihua Xiao, Fu Xiao 0001, Jianwei Zhang 0005 |
IEEE Trans. Multim. | 3 |
| 2018 | Robust image compressive sensing based on m-estimator and nonlocal low-rank regularization
Beijia Chen, Huaijiang Sun, Lei Feng 0003, Guiyu Xia, Guoqing Zhang 0002 |
Neurocomputing | 5 |
| 2018 | Mutually exclusive-KSVD: Learning a discriminative dictionary for hyperspectral image classification
Menglan Xie, Zexuan Ji, Guoqing Zhang 0002, Tao Wang 0020, Quan-Sen Sun |
Neurocomputing | 3 |
| 2018 | Multiple kernel locality-constrained collaborative representation-based discriminant projection for face recognition
Zhichao Zheng 0002, Huaijiang Sun, Guoqing Zhang 0002 |
Neurocomputing | 3 |
| 2018 | Nonlinear Low-Rank Matrix Completion for Human Motion RecoveryabstractHuman motion capture data has been widely used in many areas, but it involves a complex capture process and the captured data inevitably contains missing data due to the occlusions caused by the actor's body or clothing. Motion recovery, which aims to recover the underlying complete motion sequence from its degraded observation, still remains as a challenging task due to the nonlinear structure and kinematics property embedded in motion data. Low-rank matrix completion based methods have shown promising performance in short-time-missing motion recovery problems. However, low-rank matrix completion, which is designed for linear data, lacks the theoretic guarantee when applied to the recovery of nonlinear motion data. To overcome this drawback, we propose a tailored nonlinear matrix completion model for human motion recovery. Within the model, we first learn a combined low-rank kernel via multiple kernel learning. By exploiting the learned kernel, we embed the motion data into a high dimensional Hilbert space where motion data is of desirable low-rank and we then use the low-rank matrix completion to recover motions. In addition, we add two kinematic constraints to the proposed model to preserve the kinematics property of human motion. Extensive experiment results and comparisons with five other state-of-the-art methods demonstrate the advantage of the proposed method. Guiyu Xia, Huaijiang Sun, Beijia Chen, Qingshan Liu 0001, Lei Feng 0003, Guoqing Zhang 0002, Renlong Hang |
IEEE Trans. Image Process. | 6 |
| 2018 | Human Motion Segmentation via Robust Kernel Sparse Subspace ClusteringabstractStudies on human motion have attracted a lot of attentions. Human motion capture data, which much more precisely records human motion than videos do, has been widely used in many areas. Motion segmentation is an indispensable step for many related applications, but current segmentation methods for motion capture data do not effectively model some important characteristics of motion capture data, such as Riemannian manifold structure and containing non-Gaussian noise. In this paper, we convert the segmentation of motion capture data into a temporal subspace clustering problem. Under the framework of sparse subspace clustering, we propose to use the geodesic exponential kernel to model the Riemannian manifold structure, use correntropy to measure the reconstruction error, use the triangle constraint to guarantee temporal continuity in each cluster and use multi-view reconstruction to extract the relations between different joints. Therefore, exploiting some special characteristics of motion capture data, we propose a new segmentation method, which is robust to non-Gaussian noise, since correntropy is a localized similarity measure. We also develop an efficient optimization algorithm based on block coordinate descent method to solve the proposed model. Our optimization algorithm has a linear complexity while sparse subspace clustering is originally a quadratic problem. Extensive experiment results both on simulated noisy data set and real noisy data set demonstrate the advantage of the proposed method.Studies on human motion have attracted a lot of attentions. Human motion capture data, which much more precisely records human motion than videos do, has been widely used in many areas. Motion segmentation is an indispensable step for many related applications, but current segmentation methods for motion capture data do not effectively model some important characteristics of motion capture data, such as Riemannian manifold structure and containing non-Gaussian noise. In this paper, we convert the segmentation of motion capture data into a temporal subspace clustering problem. Under the framework of sparse subspace clustering, we propose to use the geodesic exponential kernel to model the Riemannian manifold structure, use correntropy to measure the reconstruction error, use the triangle constraint to guarantee temporal continuity in each cluster and use multi-view reconstruction to extract the relations between different joints. Therefore, exploiting some special characteristics of motion capture data, we propose a new segmentation method, which is robust to non-Gaussian noise, since correntropy is a localized similarity measure. We also develop an efficient optimization algorithm based on block coordinate descent method to solve the proposed model. Our optimization algorithm has a linear complexity while sparse subspace clustering is originally a quadratic problem. Extensive experiment results both on simulated noisy data set and real noisy data set demonstrate the advantage of the proposed method. Guiyu Xia, Huaijiang Sun, Lei Feng 0003, Guoqing Zhang 0002, Yazhou Liu |
IEEE Trans. Image Process. | 4 |
| 2017 | Dual structural consistency based multi-modal correlation propagation projections for data representation
Hongkun Ji, Quan-Sen Sun, Yun-Hao Yuan 0001, Zexuan Ji, Guoqing Zhang 0002, Lei Feng 0003 |
Multim. Tools Appl. | 5 |
| 2017 | Collaborative probabilistic labels for face recognition from single sample per person
Hongkun Ji, Quan-Sen Sun, Zexuan Ji, Yun-Hao Yuan 0001, Guoqing Zhang 0002 |
Pattern Recognit. | 5 |
| 2017 | Optimal Couple Projections for Domain Adaptive Sparse Representation-Based ClassificationabstractIn recent years, sparse representation-based classification (SRC) is one of the most successful methods and has been shown impressive performance in various classification tasks. However, when the training data have a different distribution than the testing data, the learned sparse representation may not be optimal, and the performance of SRC will be degraded significantly. To address this problem, in this paper, we propose an optimal couple projections for domain-adaptive SRC (OCPD-SRC) method, in which the discriminative features of data in the two domains are simultaneously learned with the dictionary that can succinctly represent the training and testing data in the projected space. OCPD-SRC is designed based on the decision rule of SRC, with the objective to learn coupled projection matrices and a common discriminative dictionary such that the between-class sparse reconstruction residuals of data from both domains are maximized, and the within-class sparse reconstruction residuals of data are minimized in the projected low-dimensional space. Thus, the resulting representations can well fit SRC and simultaneously have a better discriminant ability. In addition, our method can be easily extended to multiple domains and can be kernelized to deal with the nonlinear structure of data. The optimal solution for the proposed method can be efficiently obtained following the alternative optimization method. Extensive experimental results on a series of benchmark databases show that our method is better or comparable to many state-of-the-art methods. Guoqing Zhang 0002, Huaijiang Sun, Fatih Porikli, Yazhou Liu, Quan-Sen Sun |
IEEE Trans. Image Process. | 1 |
| 2016 | Learning multi-kernel multi-view canonical correlations for image recognitionabstractcanonical correlations (M 2 CCs) framework for subspace learning. In the proposed framework, the input data of each original view are mapped into multiple higher dimensional feature spaces by multiple nonlinear mappings determined by different kernels. This makes M 2 CC can discover multiple kinds of useful information of each original view in the feature spaces. With the framework, we further provide a specific multi-view feature learning method based on direct summation kernel strategy and its regularized version. The experimental results in visual recognition tasks demonstrate the effectiveness and robustness of the proposed method. Yun-Hao Yuan 0001, Yun Li 0010, Xiaobo Shen 0001, Guoqing Zhang 0002, Quan-Sen Sun |
Comput. Vis. Media | 6 |
| 2016 | Label propagation based on collaborative representation for face recognition
Guoqing Zhang 0002, Huaijiang Sun, Zexuan Ji, Quan-Sen Sun |
Neurocomputing | 1 |
| 2016 | Kernel collaborative representation based dictionary learning and discriminative projection
Guoqing Zhang 0002, Huaijiang Sun, Guiyu Xia, Quan-Sen Sun |
Neurocomputing | 1 |
| 2016 | Human motion recovery jointly utilizing statistical and kinematic information
Guiyu Xia, Huaijiang Sun, Guoqing Zhang 0002, Lei Feng 0003 |
Inf. Sci. | 3 |
| 2016 | Kernel dictionary learning based discriminant analysis
Guoqing Zhang 0002, Huaijiang Sun, Zexuan Ji, Guiyu Xia, Lei Feng 0003, Quan-Sen Sun |
J. Vis. Commun. Image Represent. | 1 |
| 2016 | Cost-sensitive dictionary learning for face recognition
Guoqing Zhang 0002, Huaijiang Sun, Zexuan Ji, Yun-Hao Yuan 0001, Quan-Sen Sun |
Pattern Recognit. | 1 |
| 2016 | Multiple Kernel Sparse Representation-Based Orthogonal Discriminative Projection and Its Cost-Sensitive ExtensionabstractSparse representation-based classification (SRC) has been developed and shown great potential for real-world application. Based on SRC, Yang et al. devised an SRC steered discriminative projection (SRC-DP) method. However, as a linear algorithm, SRC-DP cannot handle the data with highly nonlinear distribution. Kernel sparse representation-based classifier (KSRC) is a non-linear extension of SRC and can remedy the drawback of SRC. KSRC requires the use of a predetermined kernel function and selection of the kernel function and its parameters is difficult. Recently, multiple kernel learning for SRC (MKL-SRC) has been proposed to learn a kernel from a set of base kernels. However, MKL-SRC only considers the within-class reconstruction residual while ignoring the between-class relationship, when learning the kernel weights. In this paper, we propose a novel multiple kernel sparse representation-based classifier, and then we use it as a criterion to design a multiple kernel sparse representation-based orthogonal discriminative projection method. The proposed algorithm aims at learning a projection matrix and a corresponding kernel from the given base kernels such that in the low dimension subspace the between-class reconstruction residual is maximized and the within-class reconstruction residual is minimized. Furthermore, to achieve a minimum overall loss by performing recognition in the learned low-dimensional subspace, we introduce cost information into the dimensionality reduction method. The solutions for the proposed method can be efficiently found based on trace ratio optimization method. Extensive experimental results demonstrate the superiority of the proposed algorithm when compared with the state-of-the-art methods. Guoqing Zhang 0002, Huaijiang Sun, Guiyu Xia, Quan-Sen Sun |
IEEE Trans. Image Process. | 1 |