Guoqiang Liang 0001

dblp:04/10205-1 · DBLP profile ↗
← Back
29ranked-venue papers
11as first author
24since 2021 · last 2026
0000-0002-8710-5520ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 21 · 7 first-author · 17 since 2021Artificial intelligence and machine learning · 13 · 5 first-author · 11 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Attention Retention for Continual Learning with Vision Transformers
abstract
Continual learning (CL) empowers AI systems to progressively acquire knowledge from non-stationary data streams. However, catastrophic forgetting remains a critical challenge. In this work, we identify attention drift in Vision Transformers as a primary source of catastrophic forgetting, where the attention to previously learned visual concepts shifts significantly after learning new tasks. Inspired by neuroscientific insights into the selective attention in the human visual system, we propose a novel attention-retaining framework to mitigate forgetting in CL. Our method constrains attention drift by explicitly modifying gradients during backpropagation through a two-step process: 1) extracting attention maps of the previous task using a layer-wise rollout mechanism and generating instance-adaptive binary masks, and 2) when learning a new task, applying these masks to zero out gradients associated with previous attention regions, thereby preventing disruption of learned visual concepts. For compatibility with modern optimizers, the gradient masking process is further enhanced by scaling parameter updates proportionally to maintain their relative magnitudes. Experiments and visualizations demonstrate the effectiveness of our method in mitigating catastrophic forgetting and preserving visual concepts. It achieves state-of-the-art performance and exhibits robust generalizability across diverse CL scenarios.
Yue Lu 0008, Xiangyu Zhou 0002, Shizhou Zhang, Yinghui Xing, Guoqiang Liang 0001, Wencong Zhang
AAAI5
2026 Nearest-neighbor class prototype prompt and simulated logits for continual learning
Yue Lu 0008, Shizhou Zhang, Yinghui Xing, Guoqiang Liang 0001, Yanning Zhang 0001
Pattern Recognit.5
2026 Multi-Level Collaborative Distillation Meets Global Workspace Model: A Unified Framework for OCIL
abstract
Online Class-Incremental Learning (OCIL) enables models to learn continuously from non-i.i.d. data streams. Since samples of the data streams can be seen only once, it is more suitable for real-world scenarios compared to offline learning. However, this constraint intensifies the challenge for OCIL in maintaining an appropriate balance between stability and plasticity. Moreover, under stricter memory buffer constraints in real world, current replay-based methods are less effective. While ensemble methods improve plasticity, they often struggle with stability. Inspired by the Global Workspace Theory (GWT), we propose a novel approach that enhances ensemble learning through a Global Workspace Model (GWM)-a shared, implicit memory that guides the learning of multiple student models. The GWM is formed by fusing the parameters of all students within each training batch, capturing the historical learning trajectory and serving as a dynamic anchor for knowledge consolidation. Like the broadcasting mechanism of GWT, the GWM is redistributed periodically to students, stabilizing learning and promoting cross-task consistency. In addition, we introduce a multi-level collaborative distillation mechanism. It enforces peer-to-peer consistency among students and preserves historical knowledge by aligning each student with the GWM. As a result, student models remain adaptable to new tasks while maintaining previously learned knowledge, striking a better balance between stability and plasticity. Extensive experiments on three standard OCIL benchmarks show that our method delivers significant performance improvement for several OCIL models across various memory budgets. The code is available at https://github.com/susususushi/GWM.
Shibin Su, Guoqiang Liang 0001, De Cheng, Shizhou Zhang, Lingyan Ran
IEEE Trans. Image Process.2
2025 Training Consistent Mixture-of-Experts-Based Prompt Generator for Continual Learning
abstract
Visual prompt tuning-based continual learning (CL) methods have shown promising performance in exemplar-free scenarios, where their key component can be viewed as a prompt generator. Existing approaches generally rely on freezing old prompts, slow updating and task discrimination for prompt generators to preserve stability and minimize forgetting. In contrast, we introduce a novel approach that trains a consistent prompt generator to ensure stability during CL. Consistency means that for any instance from an old task, its corresponding instance-ware prompt generated by the prompt generator remains consistent even as the generator continually updates in a new task. This ensures that the representation of a specific instance remains stable across tasks and thereby prevents forgetting. We employ a mixture of experts (MoE) as the prompt generator, which contains a router and multiple experts. By deriving conditions sufficient to achieve the consistency for the MoE prompt generator, we demonstrate that: during training in a new task, if the router and experts update in the directions orthogonal to the subspaces spanned by old input features and gating vectors, respectively, the consistency can be theoretically guaranteed. To implement this orthogonality, we project parameter gradients to those orthogonal directions using the orthogonal projection matrices computed via the null space method. Extensive experiments on four class-incremental learning benchmarks validate the effectiveness and superiority of our approach.
Yue Lu 0008, Shizhou Zhang, De Cheng, Guoqiang Liang 0001, Yinghui Xing, Nannan Wang 0001, Yanning Zhang 0001
AAAI4
2025 Gradient Decomposition and Alignment for Incremental Object Detection
Wenlong Luo, Shizhou Zhang, De Cheng, Yinghui Xing, Guoqiang Liang 0001, Peng Wang 0015, Yanning Zhang 0001
ICCV5
2025 Boosting Multi-Modal Alignment: Geometric Feature Separation for Class Incremental Learning
abstract
Class Incremental Learning (CIL) aims to continually learn new classes from a stream of data without forgetting previously learned ones. Recent approaches have leveraged pre-trained models (PTMs) to improve performance, especially vision-language models, which offer better generalization than models trained solely on visual data. Many of these methods rely on simple language templates to generate class representations, which then serve as classifiers. However, due to differences between the pre-training data and downstream tasks, these textual features can become too similar for certain classes, leading to prediction errors. To address this issue, we propose a method that optimizes the geometric structure of both visual and textual features across different classes. Inspired by neural collapse theory, we introduce a multi-modal alignment strategy: for each class, a reference vector is chosen from a simplex Equiangular Tight Frame, and both the visual and textual features of the class are aligned with this vector. To better capture intra-class variations, we also construct multiple visual prototypes for each class. A multi-prototype supervised contrastive loss is then employed to pull an image feature toward the closest matching prototype of its true class and push it away from prototypes of other classes. We evaluate our approach on five widely used CIL benchmarks. The results show that our method achieves state-of-the-art performance, demonstrating its effectiveness in addressing the challenges of class incremental learning. Our code is available at https://github.com/qcNPU/NCSCMP.
Guoqiang Liang 0001, De Cheng, Shizhou Zhang, Yanning Zhang 0001
ACM Multimedia1
2025 A masking, linkage and guidance framework for online class incremental learning
Guoqiang Liang 0001, Zhaojie Chen, Shibin Su, Shizhou Zhang, Yanning Zhang 0001
Pattern Recognit.1
2025 Pseudo Labeling Methods for Semi-Supervised Semantic Segmentation: A Review and Future Perspectives
abstract
Semantic segmentation is a fundamental task in computer vision and finds extensive applications in scene understanding, medical image analysis, and remote sensing. With the advent of deep learning, significant advancements have been made in segmentation tasks. However, deep learning models require a substantial amount of labeled data for training, and accurately annotating datasets is labor-intensive and costly. Recently, numerous studies have explored the semantic segmentation task through the lens of semi-supervised learning, with the pseudo-labeling (PL) method emerging as a straightforward and widely applicable approach. This paper provides a comprehensive review and analysis of various PL methods and their applications in semi-supervised semantic segmentation (SSSS) from multiple angles. Initially, it captures the essence of individual model self-training and the collaborative training of multiple models from a model-centric viewpoint. Next, it explores strategies for refining or dismissing unreliable methods. Then, it categorizes techniques for addressing noisy PL data and inspects improvements in PL methods from the perspective of data augmentation. It further provides insights into optimization strategies. Furthermore, it examines PL methods from an application-oriented standpoint, such as in medical image segmentation and remote sensing image segmentation. Lastly, this paper evaluates the performance of cutting-edge methods on public datasets and concludes by discussing the challenges and potential directions for future research.
Lingyan Ran, Guoqiang Liang 0001, Yanning Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.3
2025 Enhancing Feature Learning With Hard Samples in Mutual Learning for Online Class Incremental Learning
abstract
Online Class-Incremental Learning (OCIL) aims to solve the problem of incrementally learning new classes from a non-i.i.d. and single-pass data stream. Compared to the offline setting, OCIL is much closer to a live learning experience requiring higher model update frequency at less computational budget. Due to its one-epoch training constraint, the model is likely to learn non-essential features and encounter the under-fitting issue, which severely affects the model's stability. In this paper, we investigate how to use hard samples to improve data variability, eventually enhancing feature learning and addressing the under-fitting problem. Specifically, by introducing a scoring function assessing the sample value, we build an OCIL formulation that simultaneously generates high-value samples and optimizes the OCIL model, improving generalization ability within the constraint of single-epoch training. Empirically, we found that strong data augmentation is a simple but effective way to generate a higher proportion of high-score samples. To make the most of these augmented samples, we design an OCIL model based on mutual learning with two networks of identical structures. Moreover, a collaborative learning mechanism is developed by aligning the features and class probabilities from the two networks to promote their interaction. Extensive experiments on three widely used datasets for OCIL have demonstrated the effectiveness of our method, obtaining superior performance to state-of-the-art methods. The code is available at https://github.com/susususushi/SDA-MCL.
Guoqiang Liang 0001, Shibin Su, De Cheng, Shizhou Zhang, Peng Wang 0015, Yanning Zhang 0001
IEEE Trans. Image Process.1
2025 Prompt-Based Modality Alignment for Effective Multi-Modal Object Re-Identification
abstract
A critical challenge for multi-modal Object Re-Identification (ReID) is the effective aggregation of complementary information to mitigate illumination issues. State-of-the-art methods typically employ complex and highly-coupled architectures, which unavoidably result in heavy computational costs. Moreover, the significant distribution gap among different image spectra hinders the joint representation of multi-modal features. In this paper, we propose a framework named as PromptMA to establish effective communication channels between different modality paths, thereby aggregating modal complementary information and bridging the distribution gap. Specifically, we inject a series of learnable multi-modal prompts into the Image Encoder and introduce a prompt exchange mechanism to enable the prompts to alternately interact with different modal token embeddings, thus capturing and distributing multi-modal features effectively. Building on top of the multi-modal prompts, we further propose Prompt-based Token Selection (PBTS) and Prompt-based Modality Fusion (PBMF) modules to achieve effective multi-modal feature fusion while minimizing background interference. Additionally, due to the flexibility of our prompt exchange mechanism, our method is well-suited to handle scenarios with missing modalities. Extensive evaluations are conducted on four widely used benchmark datasets and the experimental results demonstrate that our method achieves state-of-the-art performances, surpassing the current benchmarks by over 15% on the challenging MSVR310 dataset and by 6% on the RGBNT201. The code is available at https://github.com/FHR-L/PromptMA.
Shizhou Zhang, Wenlong Luo, De Cheng, Yinghui Xing, Guoqiang Liang 0001, Peng Wang 0015, Yanning Zhang 0001
IEEE Trans. Image Process.5
2025 Frequency-Guided Spatial Adaptation for Camouflaged Object Detection
abstract
Camouflaged object detection (COD) aims to segment camouflaged objects which exhibit very similar patterns with the surrounding environment. Recent research works have shown that enhancing the feature representation via the frequency information can greatly alleviate the ambiguity problem between the foreground objects and the background. With the emergence of vision foundation models, like InternImage, Segment Anything Model etc, adapting the pretrained model on COD tasks with a lightweight adapter module shows a novel and promising research direction. Existing adapter modules mainly care about the feature adaptation in the spatial domain. In this paper, we propose a novel frequency-guided spatial adaptation method for COD task. Specifically, we transform the input features of the adapter into frequency domain. By grouping and interacting with frequency components located within non overlapping circles in the spectrogram, different frequency components are dynamically enhanced or weakened, making the intensity of image details and contour features adaptively adjusted. At the same time, the features that are conducive to distinguishing object and background are highlighted, indirectly implying the position and shape of camouflaged object. We conduct extensive experiments on four widely adopted benchmark datasets and the proposed method outperforms 26 state-of-the-art methods with large margins. Code will be released.
Shizhou Zhang, Dexuan Kong, Yinghui Xing, Yue Lu 0008, Lingyan Ran, Guoqiang Liang 0001, Hexu Wang, Yanning Zhang 0001
IEEE Trans. Multim.6
2025 CoLeCLIP: Open-Domain Continual Learning via Joint Task Prompt and Vocabulary Learning
abstract
This article investigates the problem of continual learning (CL) of vision-language models (VLMs) in open domains, where models are required to perform continual updating and inference on a stream of datasets from diverse seen and unseen domains with novel classes. Such a capability is crucial for various applications in open environments, e.g., AI assistants, autonomous driving systems, and robotics. Current CL studies mostly focus on closed-set scenarios in a single domain with known classes. Large pretrained VLMs such as CLIP have showcased exceptional zero-shot recognition capabilities, and several recent studies have leveraged the unique characteristics of VLMs to mitigate catastrophic forgetting in CL. However, they primarily focus on closed-set CL in a single-domain dataset. Open-domain CL of large VLMs is significantly more challenging due to 1) large class correlations and domain gaps across the datasets and 2) the forgetting of zero-shot knowledge in the pretrained VLMs and the knowledge learned from the newly adapted datasets. In this work, we introduce a novel approach, termed CoLeCLIP, which learns an open-domain CL model based on CLIP. It addresses these challenges through joint learning of a set of task prompts and a cross-domain class vocabulary. Extensive experiments on 11 domain datasets show that CoLeCLIP achieves new state-of-the-art performance for open-domain CL under both task- and class-incremental learning (CIL) settings.
Guansong Pang, Wei Suo, Chenchen Jing, Yuling Xi, Lingqiao Liu, Hao Chen 0041, Guoqiang Liang 0001, Peng Wang 0015
IEEE Trans. Neural Networks Learn. Syst.8
2024 Dual Supervised Contrastive Learning Based on Perturbation Uncertainty for Online Class Incremental Learning
Shibin Su, Zhaojie Chen, Guoqiang Liang 0001, Shizhou Zhang, Yanning Zhang 0001
ICPR (9)3
2024 New Insights on Relieving Task-Recency Bias for Online Class Incremental Learning
abstract
To imitate the ability of keeping learning of human, continual learning which can learn from a never-ending data stream has attracted more interests recently. In all settings, the online class incremental learning (OCIL), where incoming samples from data stream can be used only once, is more challenging and can be encountered more frequently in real world. Actually, all continual learning models face a stability-plasticity dilemma, where the stability means the ability to preserve old knowledge while the plasticity denotes the ability to incorporate new knowledge. Although replay-based methods have shown exceptional promise, most of them concentrate on the strategy for updating and retrieving memory to keep stability at the expense of plasticity. To strike a preferable trade-off between stability and plasticity, we propose an Adaptive Focus Shifting algorithm (AFS), which dynamically adjusts focus to ambiguous samples and non-target logits in model learning. Through a deep analysis of the task-recency bias caused by class imbalance, we propose a revised focal loss to mainly keep stability. By utilizing a new weight function, the revised focal loss will pay more attention to current ambiguous samples, which are the potentially valuable samples to make model progress quickly. To promote plasticity, we introduce a virtual knowledge distillation. By designing a virtual teacher, it assigns more attention to non-target classes, which can surmount overconfidence and encourage model to focus on inter-class information. Extensive experiments on three popular datasets for OCIL have shown the effectiveness of AFS. The code will be available at https://github.com/czjghost/AFS.
Guoqiang Liang 0001, Zhaojie Chen, Zhaoqiang Chen, Shiyu Ji, Yanning Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2024 Non-Exemplar Class-Incremental Learning by Random Auxiliary Classes Augmentation and Mixed Features
abstract
Non-exemplar class-incremental learning refers to continual classifying of new and old classes without storing samples of old classes. Since only new class samples are available, catastrophic forgetting of old knowledge often occurs. In this paper, we propose an effective non-exemplar method called RAMF consisting of Random Auxiliary classes augmentation and Mixed Features. On the one hand, we design a novel random auxiliary classes augmentation method, where one augmentation is randomly selected from three augmentations and applied to inputs to generate augmented samples and extra class labels. By extending the data and label space, the model can learn more diverse and transferable representations, which can prevent the model from being biased towards learning task-specific features and facilitate the transfer among different tasks. In a word, when learning new tasks, the random auxiliary class augmentation will reduce the change of feature space and improve model generalization. On the other hand, we propose to replace the new features with mixed features for model optimization since only using new features will largely affect the previous representation embedded in the old feature space. Instead, by mixing new and old features, the cosine similarity is improved by reducing the angle between the current and old features, which allows for better stability over long-term incremental learning without increasing the computational complexity. We have conducted extensive experiments on three benchmarks CIFAR-100, Tiny-ImageNet and ImageNet-Subset, where our method outperforms the state-of-the-art non-exemplar methods and is comparable to high-performance replay-based methods.
Guoqiang Liang 0001, Zhaojie Chen, Yanning Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.2
2024 MS-DETR: Multispectral Pedestrian Detection Transformer With Loosely Coupled Fusion and Modality-Balanced Optimization
abstract
Multispectral pedestrian detection is an important task for many around-the-clock applications, since the visible and thermal modalities can provide complementary information especially under low light conditions. Due to the presence of two modalities, misalignment and modality imbalance are the most significant issues in multispectral pedestrian detection. In this paper, we propose MultiSpectral pedestrian DEtection TRansformer (MS-DETR) to fix above issues. MS-DETR consists of two modality-specific backbones and Transformer encoders, followed by a multi-modal Transformer decoder, and the visible and thermal features are fused in the multi-modal Transformer decoder. To well resist the misalignment between multi-modal images, we design a loosely coupled fusion strategy by sparsely sampling some keypoints from multi-modal features independently and fusing them with adaptively learned attention weights. Moreover, based on the insight that not only different modalities, but also different pedestrian instances tend to have different confidence scores to final detection, we further propose an instance-aware modality-balanced optimization strategy, which preserves visible and thermal decoder branches and aligns their predicted slots through an instance-wise dynamic loss. Our end-to-end MS-DETR shows superior performance on the challenging KAIST, CVC-14 and LLVIP benchmark datasets. The source code is available athttps://github.com/YinghuiXing/MS-DETR.
Yinghui Xing, Song Wang 0002, Shizhou Zhang, Guoqiang Liang 0001, Xiuwei Zhang 0001, Yanning Zhang 0001
IEEE Trans. Intell. Transp. Syst.5
2024 Dual Modality Prompt Tuning for Vision-Language Pre-Trained Model
abstract
With the emergence of large pretrained vison-language models such as CLIP, transferable representations can be adapted to a wide range of downstream tasks via prompt tuning. Prompt tuning probes for beneficial information for downstream tasks from the general knowledge stored in the pretrained model. A recently proposed method named Context Optimization (CoOp) introduces a set of learnable vectors as text prompts from the language side. However, tuning the text prompt alone can only adjust the synthesized “classifier”, while the computed visual features of the image encoder cannot be affected, thus leading to suboptimal solutions. In this article, we propose a novel dual-modality prompt tuning (DPT) paradigm through learning text and visual prompts simultaneously. To make the final image feature concentrate more on the target visual concept, a class-aware visual prompt tuning (CAVPT) scheme is further proposed in our DPT. In this scheme, the class-aware visual prompt is generated dynamically by performing the cross attention between text prompt features and image patch token embeddings to encode both the downstream task-related information and visual instance information. Extensive experimental results on 11 datasets demonstrate the effectiveness and generalization ability of the proposed method.
Yinghui Xing, Qirui Wu, De Cheng, Shizhou Zhang, Guoqiang Liang 0001, Peng Wang 0015, Yanning Zhang 0001
IEEE Trans. Multim.5
2023 Learning Conditional Attributes for Compositional Zero-Shot Learning
abstract
Compositional Zero-Shot Learning (CZSL) aims to train models to recognize novel compositional concepts based on learned concepts such as attribute-object combinations. One of the challenges is to model attributes interacted with different objects, e.g., the attribute “wet” in “wet apple” and “wet cat” is different. As a solution, we provide analysis and argue that attributes are conditioned on the recognized object and input image and explore learning conditional attribute embeddings by a proposed attribute learning framework containing an attribute hyper learner and an attribute base learner. By encoding conditional attributes, our model enables to generate flexible attribute embeddings for generalization from seen to unseen compositions. Experiments on CZSL benchmarks, including the more challenging C-GQA dataset, demonstrate better performances compared with other state-of-the-art approaches and validate the importance of learning conditional attributes. Code‡1Gllee:https://gitee.com/wqshmzh/canet-czsl is available at https://github.com/wqshmzh/CANet-CZSL.
Qingsheng Wang, Lingqiao Liu, Chenchen Jing, Hao Chen 0041, Guoqiang Liang 0001, Peng Wang 0015, Chunhua Shen
CVPR5
2023 Ground-to-Aerial Person Search: Benchmark Dataset and Approach
abstract
In this work, we construct a large-scale dataset for Ground-to-Aerial Person Search, named G2APS, which contains 31,770 images of 260,559 annotated bounding boxes for 2,644 identities appearing in both of the UAVs and ground surveillance cameras. To our knowledge, this is the first dataset for cross-platform intelligent surveillance applications, where the UAVs could work as a powerful complement for the ground surveillance cameras. To more realistically simulate the actual cross-platform Ground-to-Aerial surveillance scenarios, the surveillance cameras are fixed about 2 meters above the ground, while the UAVs capture videos of persons at different location, with a variety of view-angles, flight attitudes and flight modes. Therefore, the dataset has the following unique characteristics: 1) drastic view-angle changes between query and gallery person images from cross-platform cameras; 2) diverse resolutions, poses and views of the person images under 9 rich real-world scenarios. On basis of the G2APS benchmark dataset, we demonstrate detailed analysis about current two-step and end-to-end person search methods, and further propose a simple yet effective knowledge distillation scheme on the head of the ReID network, which achieves state-of-the-art performances on both of the G2APS and the previous two public person search datasets, i.e., PRW and CUHK-SYSU. The dataset and source code available on https://github.com/yqc123456/HKD_for_person_search.
Shizhou Zhang, Qingchun Yang, De Cheng, Yinghui Xing, Guoqiang Liang 0001, Peng Wang 0015, Yanning Zhang 0001
ACM Multimedia5
2022 Video summarization with a dual-path attentive network
Guoqiang Liang 0001, Yanbing Lv, Shucheng Li, Xiahong Wang, Yanning Zhang 0001
Neurocomputing1
2022 Video summarization with a convolutional attentive adversarial network
Guoqiang Liang 0001, Yanbing Lv, Shucheng Li, Shizhou Zhang, Yanning Zhang 0001
Pattern Recognit.1
2021 Local-enhanced Interaction for Temporal Moment Localization
abstract
Temporal moment localization via language aims to localize a video span in an untrimmed video which best matches the given natural language query. In most previous works, they try to match the whole query feature with multiple moment proposals, or match a global video embedding with phrase or word level query features. However, these coarse interaction models will become insufficient when the query-video contains more complex relationship. To address this issue, we propose a multi-branches interaction model for temporal moment localization. Specifically, the query sentence and video are encoded into multiple feature embeddings over several semantic sub-spaces. Then, each phrase embedding filters on a video feature to generate an attention sequence, which is used to re-weight the video features. Moreover, a dynamic pointer decoder is developed to iteratively regress the temporal boundary, which can prevent our model from falling into a local optimum. To validate the proposed method, we have conducted extensive experiments on two popular benchmark datasets Charade-STA and TACoS. The experimental performance surpasses other state-of-the-arts methods, which demonstrates the effectiveness of our proposed model.
Guoqiang Liang 0001, Shiyu Ji, Yanning Zhang 0001
ICMR1
2021 An adversarial human pose estimation network injected with graph structure
Peng Wang 0015, Guoqiang Liang 0001, Chunhua Shen
Pattern Recognit.3
2021 Attend to the Difference: Cross-Modality Person Re-Identification via Contrastive Correlation
abstract
The problem of cross-modality person re-identification has been receiving increasing attention recently, due to its practical significance. Motivated by the fact that human usually attend to the difference when they compare two similar objects, we propose a dual-path cross-modality feature learning framework which preserves intrinsic spatial structures and attends to the difference of input cross-modality image pairs. Our framework is composed by two main components: a Dual-path Spatial-structure-preserving Common Space Network (DSCSN) and a Contrastive Correlation Network (CCN). The former embeds cross-modality images into a common 3D tensor space without losing spatial structures, while the latter extracts contrastive features by dynamically comparing input image pairs. Note that the representations generated for the input RGB and Infrared images are mutually dependant to each other. We conduct extensive experiments on two public available RGB-IR ReID datasets, SYSU-MM01 and RegDB, and our proposed method outperforms state-of-the-art algorithms by a large margin with both full and simplified evaluation modes.
Shizhou Zhang, Peng Wang 0015, Guoqiang Liang 0001, Xiuwei Zhang 0001, Yanning Zhang 0001
IEEE Trans. Image Process.4
2019 Cross-View Person Identification Based on Confidence-Weighted Human Pose Matching
abstract
Cross-view person identification (CVPI) from multiple temporally synchronized videos taken by multiple wearable cameras from different, varying views is a very challenging but important problem, which has attracted more interest recently. Current state-of-the-art performance of CVPI is achieved by matching appearance and motion features across videos, while the matching of pose features does not work effectively given the high inaccuracy of the 3D pose estimation on videos/images collected in the wild. To address this problem, we first introduce a new metric of confidence to the estimated location of each human-body joint in 3D human pose estimation. Then, a mapping function, which can be hand-crafted or learned directly from the datasets, is proposed to combine the inaccurately estimated human pose and the inferred confidence metric to accomplish CVPI. Specifically, the joints with higher confidence are weighted more in the pose matching for CVPI. Finally, the estimated pose information is integrated into the appearance and motion features to boost the CVPI performance. In the experiments, we evaluate the proposed method on three wearable-camera video datasets and compare the performance against several other existing CVPI methods. The experimental results show the effectiveness of the proposed confidence metric, and the integration of pose, appearance, and motion produces a new state-of-the-art CVPI performance.
Guoqiang Liang 0001, Xuguang Lan, Xingyu Chen 0001, Song Wang 0002, Nanning Zheng 0001
IEEE Trans. Image Process.1
2018 Cross-View Person Identification by Matching Human Poses Estimated With Confidence on Each Body Joint
abstract
Cross-view person identification (CVPI) from multiple temporally synchronized videos taken by multiple wearable cameras from different, varying views is a very challenging but important problem, which has attracted more interests recently. Current state-of-the-art performance of CVPI is achieved by matching appearance and motion features across videos, while the matching of pose features does not work effectively given the high inaccuracy of the 3D human pose estimation on videos/images collected in the wild. In this paper, we introduce a new metric of confidence to the 3D human pose estimation and show that the combination of the inaccurately estimated human pose and the inferred confidence metric can be used to boost the CVPI performance---the estimated pose information can be integrated to the appearance and motion features to achieve the new state-of-the-art CVPI performance. More specifically, the estimated confidence metric is measured at each human-body joint and the joints with higher confidence are weighted more in the pose matching for CVPI. In the experiments, we validate the proposed method on three wearable-camera video datasets and compare the performance against several other existing CVPI methods.
Guoqiang Liang 0001, Xuguang Lan, Song Wang 0002, Nanning Zheng 0001
AAAI1
2018 A Limb-Based Graphical Model for Human Pose Estimation
abstract
Modeling the relationship among human joints is one of the most important components in human pose estimation. Most of previous methods define this relationship as a geometric constraint on the relative locations of two neighboring joints. In this constraint, the local appearance of the region connecting two neighboring joints is ignored. However, discarding this image appearance leads to some severe problems, such as double-counting and localization failure when the human pose is rare in the training dataset. Moreover, this image appearance, called human limb, plays an important role in human pose estimation in human visual system. Due to these reasons, we propose to solve a new task: human limb detection, which aims at detecting and representing this local image appearance. We combine this task with human joint localization as a unified framework. After getting the initial detections, we design a two-steps graphical model to capture the spatial relationship among human joints and limbs in a coarse to fine way. We evaluate the proposed method on two widely used datasets for human pose estimation: 1) frame labeled in cinema and 2) leeds sports pose datasets. The experiments results show the effectiveness of our method.
Guoqiang Liang 0001, Xuguang Lan, Jiang Wang 0001, Jianji Wang 0001, Nanning Zheng 0001
IEEE Trans. Syst. Man Cybern. Syst.1
2017 Pose-and-illumination-invariant face representation via a triplet-loss trained deep reconstruction model
Xingyu Chen 0001, Xuguang Lan, Guoqiang Liang 0001, Nanning Zheng 0001
Multim. Tools Appl.3
2016 Human pose estimation based on human limbs
abstract
Modeling the relationship among human joints is one of the most important components in human pose estimation. Previous methods usually define this relationship as geometric constraints on the relative location of two neighboring joints. In this definition, the local image appearance of the region connecting two neighboring joints is ignored. In fact, this image appearance, called human limb, plays an important role in human joint localization in human visual system. To make full use of this local image appearance, we propose to solve a new task: human limb detection. We combine it with human joint localization in one deep convolutional neural network. After getting coarse results, we employ a graphical model to remove false positive detections. Besides, shallow and deep features are combined in this model. We evaluate our method on the FLIC and LSP datasets. The experiments results show the effectiveness of our method.
Guoqiang Liang 0001, Xuguang Lan, Jiang Wang 0001, Nanning Zheng 0001
ICPR1