Pingyang Dai

dblp:04/8207 · DBLP profile ↗
← Back
31ranked-venue papers
4as first author
26since 2021 · last 2026
0000-0001-9780-271XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 20 · 4 first-author · 17 since 2021Artificial intelligence and machine learning · 18 · 1 first-author · 17 since 2021Human-computer interaction and ubiquitous computing · 2Applied, interdisciplinary, general and emerging computing · 2
YearPublicationVenuePosition
2026 RIS-LAD: A Benchmark and Model for Referring Image Segmentation in Low-Altitude Drone Imagery
abstract
Referring Image Segmentation (RIS), which aims to segment specific objects based on natural language descriptions, plays an essential role in vision-language understanding. Despite its progress in remote sensing applications, RIS under Low-Altitude Drone (LAD) scenarios remains underexplored, as existing datasets and methods are typically designed for high-altitude and static-view imagery. They struggled to handle the unique characteristics of LAD views, such as diverse viewpoints and high object density. In this paper, we propose RIS-LAD, the first fine-grained RIS benchmark tailored for LAD scenarios, featuring 13,871 meticulously annotated image-text-mask triplets collected from real-world drone footage with emphasis on small, densely cluttered objects and multi-view perspectives. Additionally, we propose the Semantic-Aware Adaptive Reasoning Network, which decomposes and adaptively routes semantic information to different network stages rather than uniformly injecting all linguistic features. Specifically, the Category-Dominated Linguistic Enhancement aligns visual features with object categories during early encoding, while the Adaptive Reasoning Fusion Module dynamically selects semantic cues across scales to enhance reasoning in complex scenes. Extensive experiments reveal that RIS-LAD presents substantial challenges to state-of-the-art RIS algorithms, and also demonstrate the effectiveness of our proposed model in addressing these challenges.
YingShi Luan, Zhudi Chen, Guangyue Meng, Pingyang Dai, Liujuan Cao
AAAI5
2026 Knowing Where to Focus: Attention-Guided Alignment for Text-based Person Search
Pingyang Dai, Jie Chen 0001, Liujuan Cao, Rongrong Ji
Int. J. Comput. Vis.3
2025 DS-VLM: Diffusion Supervision Vision Language Model
abstract
Vision-Language Models (VLMs) face two critical limitations in visual representation learning: degraded supervision due to information loss during gradient propagation, and the inherent semantic sparsity of textual supervision compared to visual data. We propose the Diffusion Supervision Vision-Language Model (DS-VLM), a plug-and-play framework that introduces diffusion-based direct supervision for vision-language alignment. By reconstructing input images through a diffusion model conditioned on outputs of the visual encoder and the connector, our method establishes a short-path gradient propagation channel from pixel space to visual features. This approach simultaneously preserves high-level semantic alignment through conventional text supervision while enhancing visual feature quality via pixel-level reconstruction constraints. Extensive experiments conducted across various visual encoders and LLMs of different scales demonstrate the effectiveness of our approach.
Yunhang Shen, Jie Li 0052, Xing Sun 0001, Pingyang Dai, Liujuan Cao, Rongrong Ji
ICML5
2025 FlexiReID: Adaptive Mixture of Expert for Multi-Modal Person Re-Identification
abstract
Multimodal person re-identification (Re-ID) aims to match pedestrian images across different modalities. However, most existing methods focus on limited cross-modal settings and fail to support arbitrary query-retrieval combinations, hindering practical deployment. We propose FlexiReID, a flexible framework that supports seven retrieval modes across four modalities: RGB, infrared, sketches, and text. FlexiReID introduces an adaptive mixture-of-experts (MoE) mechanism to dynamically integrate diverse modality features and a cross-modal query fusion module to enhance multimodal feature extraction. To facilitate comprehensive evaluation, we construct CIRS-PEDES, a unified dataset extending four popular Re-ID datasets to include all four modalities. Extensive experiments demonstrate that FlexiReID achieves state-of-the-art performance and offers strong generalization in complex scenarios.
Yunhang Shen, Chengmao Cai, Xing Sun 0001, Pingyang Dai, Liujuan Cao, Rongrong Ji
ICML6
2025 CMAN: Compact Modality Alignment Network With Dual Stream Transformer For Visible-Infrared Person Re-identification
abstract
Visible-infrared person re-identification (VI-ReID) aims to match pedestrian images captured by visible and infrared cameras. The main challenge lies in the severe cross-modality and intra-modality differences between visible light (VIS) and infrared (IR) images. Most current CNN-based methods achieve cross-modality retrieval by mapping two images into a high-dimensional subspace and exploiting modality-shared features, making it difficult to mine diverse cross-modality representations and effectively capture global image dependencies. To address these issues, we propose a compact modality alignment network (CMAN) based on the dual stream ViT architecture to explore a novel modality alignment approach for VI-ReID. Specifically, we first deploy a dual deep network based on vision transformer and divide the layers of the self-attention mechanism into heterogeneous and isomorphic modules, allowing us to extract modality-specific features and shared features of each image at different stages. Then, in the heterogeneous module, we use multiple class tokens in each modality to represent multiple embedding spaces and apply a Dynamic Controller (DC) in these spaces to push each class token away from each other by adaptively adjusting the weight of each class token, which makes each embedding space diverse and compact, thereby improving the discrimination of modality-specific features in each modality. Finally, in the isomorphic part, we use the Token Permutation (TP) module to permute and concatenate class tokens in different modalities. Not only helps align shared features of modalities, it also allows the class token in the current modality to perceive the local details of another modality. Extensive experiments on the public SYSU-MM01, RegDB, and LLCM datasets demonstrate the superiority of the proposed CMAN over state-of-the-art methods.
Xuyang Song, Pingyang Dai
IJCNN2
2025 GPT-ReID: Learning Fine-grained Representation with GPT for Text-based Person Retrieval
abstract
Text-based Person Retrieval (TBPR) is a challenging task that aims to retrieve pedestrian images according to natural language descriptions. Existing works mainly focus on discriminative feature learning via exploring cross-modal matching methods, while the overfitting issues caused by insufficient labeled data and the absence of well-designed auxiliary tasks are often overlooked. Motivated by the recent progress of large language models (LLMs), we propose a novel method named GPT-ReID for TBPR, which aims to leverage the strong comprehension of LLMs to alleviate the overfitting risk. Specifically, based on the great power of GPT, GPT-ReID first introduces an adversarial text generation scheme called GPTGAN, which aims to generate comprehensive strong positive captions and deceptive hard negative captions through the original captions for a single image. Furthermore, a joint auxiliary learning strategy is also proposed which contains Multi-Relation Aware (MRA), Keywords Masked Language Model (KMLM), and Keywords Replacement Detection (KRD), to facilitate global- and token-level optimization, enhancing cross-modal granular representation alignment. Extensive experiments on a set of highly competitive benchmark datasets validate the merits of the proposed GPT-ReID against a flurry of state-of-the-art methods, with Rank-1 accuracy reaching 78.42%, 69.43%, and 70.06% on CUHK-PEDS, ICFG-PEDS, and RSTPReid, respectively.
Pingyang Dai, Liujuan Cao, Rongrong Ji
ACM Multimedia3
2025 Unsupervised Domain Adaptation on Person Reidentification Via Dual-Level Asymmetric Mutual Learning
abstract
Unsupervised domain adaptation (UDA) person reidentification (Re-ID) aims to identify pedestrian images within an unlabeled target domain with an auxiliary labeled source-domain dataset. Many existing works attempt to recover reliable identity information by considering multiple homogeneous networks. And take these generated labels to train the model in the target domain. However, these homogeneous networks identify people in approximate subspaces and equally exchange their knowledge with others or their mean net to improve their ability, inevitably limiting the scope of available knowledge and putting them into the same mistake. This article proposes a dual-level asymmetric mutual learning (DAML) method to learn discriminative representations from a broader knowledge scope with diverse embedding spaces. Specifically, two heterogeneous networks mutually learn knowledge from asymmetric subspaces through the pseudo label generation in a hard distillation manner. The knowledge transfer between two networks is based on an asymmetric mutual learning (AML) manner. The teacher network learns to identify both the target and source domain while adapting to the target domain distribution based on the knowledge of the student. Meanwhile, the student network is trained on the target dataset and employs the ground-truth label through the knowledge of the teacher. Extensive experiments in Market-1501, CUHK-SYSU, and MSMT17 public datasets verified the superiority of DAML over state-of-the-arts (SOTA).
Qiong Wu 0012, Jiahan Li, Pingyang Dai, Qixiang Ye, Liujuan Cao, Yongjian Wu 0001, Rongrong Ji
IEEE Trans. Neural Networks Learn. Syst.3
2025 CycleTrans: Learning Neutral Yet Discriminative Features via Cycle Construction for Visible- Infrared Person Re-Identification
abstract
Visible-infrared person re-identification (VI-ReID) is the task of matching the same individuals across the visible and infrared modalities. Its main challenge lies in the modality gap caused by the cameras operating on different spectra. Existing VI-ReID methods mainly focus on learning general features across modalities, often at the expense of feature discriminability. To address this issue, we present a novel cycle-construction-based network for neutral yet discriminative feature learning, termed CycleTrans. Specifically, CycleTrans uses a lightweight knowledge capturing module (KCM) to capture rich semantics from the modality-relevant feature maps according to pseudo anchors. Afterward, a discrepancy modeling module (DMM) is deployed to transform these features into neutral ones according to the modality-irrelevant prototypes. To ensure feature discriminability, another two KCMs are further deployed for feature cycle constructions. With cycle construction, our method can learn effective neutral features for visible and infrared images while preserving their salient semantics. Extensive experiments on SYSU-MM01 and RegDB datasets validate the merits of CycleTrans against a flurry of state-of-the-art (SOTA) methods, on rank-1 in SYSU-MM01 and on rank-1 in RegDB. Our code is available at https://github.com/DoubtedSteam/CycleTrans.
Qiong Wu 0012, Jiaer Xia, Pingyang Dai, Yiyi Zhou, Yongjian Wu 0001, Rongrong Ji
IEEE Trans. Neural Networks Learn. Syst.3
2024 Semi-Supervised Blind Image Quality Assessment through Knowledge Distillation and Incremental Learning
abstract
Blind Image Quality Assessment (BIQA) aims to simulate human assessment of image quality. It has a great demand for labeled data, which is often insufficient in practice. Some researchers employ unsupervised methods to address this issue, which is challenging to emulate the human subjective system. To this end, we introduce a unified framework that combines semi-supervised and incremental learning to address the mentioned issue. Specifically, when training data is limited, semi-supervised learning is necessary to infer extensive unlabeled data. To facilitate semi-supervised learning, we use knowledge distillation to assign pseudo-labels to unlabeled data, preserving analytical capability. To gradually improve the quality of pseudo labels, we introduce incremental learning. However, incremental learning can lead to catastrophic forgetting. We employ Experience Replay by selecting representative samples during multiple rounds of semi-supervised learning, to alleviate forgetting and ensure model stability. Experimental results show that the proposed approach achieves state-of-the-art performance across various benchmark datasets. After being trained on the LIVE dataset, our method can be directly transferred to the CSIQ dataset. Compared with other methods, it significantly outperforms unsupervised methods on the CSIQ dataset with a marginal performance drop (-0.002) on the LIVE dataset. In conclusion, our proposed method demonstrates its potential to tackle the challenges in real-world production processes.
Wensheng Pan, Timin Gao, Yan Zhang 0109, Xiawu Zheng, Yunhang Shen, Ke Li 0015, Runze Hu, Yutao Liu 0002, Pingyang Dai
AAAI9
2024 Occluded Person Re-identification via Saliency-Guided Patch Transfer
abstract
While generic person re-identification has made remarkable improvement in recent years, these methods are designed under the assumption that the entire body of the person is available. This assumption brings about a significant performance degradation when suffering from occlusion caused by various obstacles in real-world applications. To address this issue, data-driven strategies have emerged to enhance the model's robustness to occlusion. Following the random erasing paradigm, these strategies typically employ randomly generated noise to supersede randomly selected image regions to simulate obstacles. However, the random strategy is not sensitive to location and content, meaning they cannot mimic real-world occlusion cases in application scenarios. To overcome this limitation and fully exploit the real scene information in datasets, this paper proposes a more intuitive and effective data-driven strategy named Saliency-Guided Patch Transfer (SPT). Combined with the vision transformer, SPT divides person instances and background obstacles using salient patch selection. By transferring person instances to different background obstacles, SPT can easily generate photo-realistic occluded samples. Furthermore, we propose an occlusion-aware Intersection over Union (OIoU) with mask-rolling to filter the more suitable combination and a class-ignoring strategy to achieve more stable processing. Extensive experimental evaluations conducted on occluded and holistic person re-identification benchmarks demonstrate that SPT provides a significant performance gain among different ViT-based ReID algorithms on occluded ReID.
Jiaer Xia, Pingyang Dai, Yongjian Wu 0001, Liujuan Cao
AAAI4
2024 Attention Disturbance and Dual-Path Constraint Network for Occluded Person Re-identification
abstract
Occluded person re-identification (Re-ID) aims to address the potential occlusion problem when matching occluded or holistic pedestrians from different camera views. Many methods use the background as artificial occlusion and rely on attention networks to exclude noisy interference. However, the significant discrepancy between simple background occlusion and realistic occlusion can negatively impact the generalization of the network. To address this issue, we propose a novel transformer-based Attention Disturbance and Dual-Path Constraint Network (ADP) to enhance the generalization of attention networks. Firstly, to imitate real-world obstacles, we introduce an Attention Disturbance Mask (ADM) module that generates an offensive noise, which can distract attention like a realistic occluder, as a more complex form of occlusion. Secondly, to fully exploit these complex occluded images, we develop a DualPath Constraint Module (DPC) that can obtain preferable supervision information from holistic images through dualpath interaction. With our proposed method, the network can effectively circumvent a wide variety of occlusions using the basic ViT baseline. Comprehensive experimental evaluations conducted on person re-ID benchmarks demonstrate the superiority of ADP over state-of-the-art methods.
Jiaer Xia, Pingyang Dai, Ming-Bo Zhao, Yongjian Wu 0001, Liujuan Cao
AAAI3
2024 TF-FAS: Twofold-Element Fine-Grained Semantic Guidance for Generalizable Face Anti-spoofing
Ke-Yue Zhang, Taiping Yao, Qianyu Zhou 0001, Shouhong Ding, Pingyang Dai, Rongrong Ji
ECCV (7)6
2024 Adaptive Feature Selection for No-Reference Image Quality Assessment by Mitigating Semantic Noise Sensitivity
abstract
The current state-of-the-art No-Reference Image Quality Assessment (NR-IQA) methods typically rely on feature extraction from upstream semantic backbone networks, assuming that all extracted features are relevant. However, we make a key observation that not all features are beneficial, and some may even be harmful, necessitating careful selection. Empirically, we find that many image pairs with small feature spatial distances can have vastly different quality scores, indicating that the extracted features may contain quality-irrelevant noise. To address this issue, we propose a Quality-Aware Feature Matching IQA Metric (QFM-IQM) that employs an adversarial perspective to remove harmful semantic noise features from the upstream task. Specifically, QFM-IQM enhances the semantic noise distinguish capabilities by matching image pairs with similar quality scores but varying semantic features as adversarial semantic noise and adaptively adjusting the upstream task’s features by reducing sensitivity to adversarial noise perturbation. Furthermore, we utilize a distillation framework to expand the dataset and improve the model’s generalization ability. Extensive experiments conducted on eight standard IQA datasets have demonstrated the effectiveness of our proposed QFM-IQM.
Timin Gao, Runze Hu, Yan Zhang 0109, Shengchuan Zhang, Xiawu Zheng, Jingyuan Zheng, Yunhang Shen, Ke Li 0015, Yutao Liu 0002, Pingyang Dai, Rongrong Ji
ICML11
2024 Integrating Global Context Contrast and Local Sensitivity for Blind Image Quality Assessment
abstract
Blind Image Quality Assessment (BIQA) mirrors subjective made by human observers. Generally, humans favor comparing relative qualities over predicting absolute qualities directly. However, current BIQA models focus on mining the "local" context, i.e., the relationship between information among individual images and the absolute quality of the image, ignoring the "global" context of the relative quality contrast among different images in the training data. In this paper, we present the Perceptual Context and Sensitivity BIQA (CSIQA), a novel contrastive learning paradigm that seamlessly integrates "global” and "local” perspectives into the BIQA. Specifically, the CSIQA comprises two primary components: 1) A Quality Context Contrastive Learning module, which is equipped with different contrastive learning strategies to effectively capture potential quality correlations in the global context of the dataset. 2) A Quality-aware Mask Attention Module, which employs the random mask to ensure the consistency with visual local sensitivity, thereby improving the model’s perception of local distortions. Extensive experiments on eight standard BIQA datasets demonstrate the superior performance to the state-of-the-art BIQA methods.
Runze Hu, Jingyuan Zheng, Yan Zhang 0109, Shengchuan Zhang, Xiawu Zheng, Ke Li 0015, Yunhang Shen, Yutao Liu 0002, Pingyang Dai, Rongrong Ji
ICML10
2024 RLE: A Unified Perspective of Data Augmentation for Cross-Spectral Re-Identification
abstract
This paper makes a step towards modeling the modality discrepancy in the cross-spectral re-identification task. Based on the Lambertain model, we observe that the non-linear modality discrepancy mainly comes from diverse linear transformations acting on the surface of different materials. From this view, we unify all data augmentation strategies for cross-spectral re-identification as mimicking such local linear transformations and categorize them into moderate transformation and radical transformation. By extending the observation, we propose a Random Linear Enhancement (RLE) strategy which includes Moderate Random Linear Enhancement (MRLE) and Radical Random Linear Enhancement (RRLE) to push the boundaries of both types of transformation. Moderate Random Linear Enhancement is designed to provide diverse image transformations that satisfy the original linear correlations under constrained conditions, whereas Radical Random Linear Enhancement seeks to generate local linear transformations directly without relying on external information. The experimental results not only demonstrate the superiority and effectiveness of RLE but also confirm its great potential as a general-purpose data augmentation for cross-spectral re-identification.
Keke Han, Pingyang Dai, Yan Zhang 0109, Yongjian Wu 0001, Rongrong Ji
NeurIPS4
2024 CPE COIN++: Towards Optimized Implicit Neural Representation Compression Via Chebyshev Positional Encoding
Haocheng Chu, Shaohui Dai, Wenqi Ding, Tianshuo Xu, Pingyang Dai, Shengchuan Zhang, Yan Zhang 0109, Xiang Chang, Chih-Min Lin, Fei Chao 0001, Changjiang Shang, Qiang Shen 0001
PRCV (9)6
2023 Learning Occlusion Disentanglement with Fine-grained Localization for Occluded Person Re-identification
abstract
Person re-identification (Re-ID) has been extensively investigated in recent years. However, many existing paradigms rely on holistic person regions for matching, disregarding the challenges posed by occlusions in real-world scenarios. Recent methods have explored occlusion augmentation or external semantic cues. Nevertheless, these approaches tend to be coarse-grained, discarding valuable semantic information in local regions when determining them as occlusions. In this paper, we propose a Fine-grained Occlusion Disentanglement Network (FODN) that can extract more information from limited person regions. Specifically, we propose a fine-grained occlusion augmentation scheme to generate diverse occlusion data and employ bilinear interpolation and downsampling strategies to obtain fine-grained occlusion labels. We then design an occlusion feature disentanglement Module that decouples norm and angle from features and supervises the occlusion-aware task using the aforementioned occlusion labeling and person re-identification tasks, respectively, resulting in more robust features. Additionally, we propose a dynamic local weight controller to balance the relative importance of various human body parts, thereby improving the model's ability to mine more effective local features from limited human body regions after occlusion removal. Comprehensive experiments on various person Re-ID benchmarks demonstrate the superiority of FODN over state-of-the-art methods.
Yan Zhang 0109, Pingyang Dai, Yongjian Wu 0001, Rongrong Ji
ACM Multimedia5
2023 Cross-Dataset Distillation with Multi-tokens for Image Quality Assessment
Timin Gao, Weixuan Jin, Bokai Lai, Runze Hu, Yan Zhang 0109, Pingyang Dai
PRCV (6)7
2023 Quality-Aware CLIP for Blind Image Quality Assessment
Wensheng Pan, Zhifu Yang, DingMing Liu, Chenxin Fang, Yan Zhang 0109, Pingyang Dai
PRCV (6)6
2023 Enhancing Text-Image Person Retrieval Through Nuances Varied Sample
Jiaer Xia, Haozhe Yang, Yan Zhang 0002, Pingyang Dai
PRCV (1)4
2023 Prompt Based Lifelong Person Re-identification
Chengde Yang, Yan Zhang 0109, Pingyang Dai
PRCV (12)3
2022 Dynamic Prototype Mask for Occluded Person Re-Identification
abstract
Although person re-identification has achieved an impressive improvement in recent years, the common occlusion case caused by different obstacles is still an unsettled issue in real application scenarios. Existing methods mainly address this issue by employing body clues provided by an extra network to distinguish the visible part. Nevertheless, the inevitable domain gap between the assistant model and the ReID datasets has highly increased the difficulty to obtain an effective and efficient model. To escape from the extra pre-trained networks and achieve an automatic alignment in an end-to-end trainable network, we propose a novel Dynamic Prototype Mask (DPM) based on two self-evident prior knowledge. Specifically, we first devise a Hierarchical Mask Generator which utilizes the hierarchical semantic to select the visible pattern space between the high-quality holistic prototype and the feature representation of the occluded input image. Under this condition, the occluded representation could be well aligned in a selected subspace spontaneously. Then, to enrich the feature representation of the high-quality holistic prototype and provide a more complete feature space, we introduce a Head Enrich Module to encourage different heads to aggregate different patterns representation in the whole image. Extensive experimental evaluations conducted on occluded and holistic person re-identification benchmarks demonstrate the superior performance of the DPM over the state-of-the-art methods.
Pingyang Dai, Rongrong Ji, Yongjian Wu 0001
ACM Multimedia2
2022 Disentangling Task-Oriented Representations for Unsupervised Domain Adaptation
abstract
Unsupervised domain adaptation (UDA) aims to address the domain-shift problem between a labeled source domain and an unlabeled target domain. Many efforts have been made to eliminate the mismatch between the distributions of training and testing data by learning domain-invariant representations. However, the learned representations are usually not task-oriented, i.e., being class-discriminative and domain-transferable simultaneously. This drawback limits the flexibility of UDA in complicated open-set tasks where no labels are shared between domains. In this paper, we break the concept of task-orientation into task-relevance and task-irrelevance, and propose a dynamic task-oriented disentangling network (DTDN) to learn disentangled representations in an end-to-end fashion for UDA. The dynamic disentangling network effectively disentangles data representations into two components: the task-relevant ones embedding critical information associated with the task across domains, and the task-irrelevant ones with the remaining non-transferable or disturbing information. These two components are regularized by a group of task-specific objective functions across domains. Such regularization explicitly encourages disentangling and avoids the use of generative models or decoders. Experiments in complicated, open-set scenarios (retrieval tasks) and empirical benchmarks (classification tasks) demonstrate that the proposed method captures rich disentangled information and achieves superior performance.
Pingyang Dai, Peixian Chen, Qiong Wu 0012, Xiaopeng Hong, Qixiang Ye, Qi Tian 0001, Chia-Wen Lin, Rongrong Ji
IEEE Trans. Image Process.1
2021 Dual Distribution Alignment Network for Generalizable Person Re-Identification
abstract
Domain generalization (DG) offers a preferable real-world setting for Person Re-Identification (Re-ID), which trains a model using multiple source domain datasets and expects it to perform well in an unseen target domain without any model updating. Unfortunately, most DG approaches are designed explicitly for classification tasks, which fundamentally differs from the retrieval task Re-ID. Moreover, existing applications of DG in Re-ID cannot correctly handle the massive variation among Re-ID datasets. In this paper, we identify two fundamental challenges in DG for Person Re-ID: domain-wise variations and identity-wise similarities. To this end, we propose an end-to-end Dual Distribution Alignment Network (DDAN) to learn domain-invariant features with dual-level constraints: the domain-wise adversarial feature learning and the identity-wise similarity enhancement. These constraints effectively reduce the domain-shift among multiple source domains further while agreeing to real-world scenarios. We evaluate our method in a large-scale DG Re-ID benchmark and compare it with various cutting-edge DG approaches. Quantitative results show that DDAN achieves state-of-the-art performance.
Peixian Chen, Pingyang Dai, Jianzhuang Liu, Feng Zheng 0001, Mingliang Xu 0001, Qi Tian 0001, Rongrong Ji
AAAI2
2021 Discover Cross-Modality Nuances for Visible-Infrared Person Re-Identification
abstract
Visible-infrared person re-identification (Re-ID) aims to match the pedestrian images of the same identity from different modalities. Existing works mainly focus on alleviating the modality discrepancy by aligning the distributions of features from different modalities. However, nuanced but discriminative information, such as glasses, shoes, and the length of clothes, has not been fully explored, especially in the infrared modality. Without discovering nuances, it is challenging to match pedestrians across modalities using modality alignment solely, which inevitably reduces feature distinctiveness. In this paper, we propose a joint Modality and Pattern Alignment Network (MPANet) to discover cross-modality nuances in different patterns for visible-infrared person Re-ID, which introduces a modality alleviation module and a pattern alignment module to jointly extract discriminative features. Specifically, we first propose a modality alleviation module to dislodge the modality information from the extracted feature maps. Then, We devise a pattern alignment module, which generates multiple pattern maps for the diverse patterns of a person, to discover nuances. Finally, we introduce a mutual mean learning fashion to alleviate the modality discrepancy and propose a center cluster loss to guide both identity learning and nuances discovering. Extensive experiments on the public SYSU-MM01 and RegDB datasets demonstrate the superiority of MPANet over state-of-the-arts.
Qiong Wu 0012, Pingyang Dai, Jie Chen 0001, Chia-Wen Lin, Yongjian Wu 0001, Feiyue Huang, Bineng Zhong 0001, Rongrong Ji
CVPR2
2021 Occlude Them All: Occlusion-Aware Attention Network for Occluded Person Re-ID
abstract
Person Re-Identification (ReID) has achieved remarkable performance along with the deep learning era. However, most approaches carry out ReID only based upon holistic pedestrian regions. In contrast, real-world scenarios involve occluded pedestrians, which provide partial visual appearances and destroy the ReID accuracy. A common strategy is to locate visible body parts by auxiliary model, which however suffers from significant domain gaps and data bias issues. To avoid such problematic models in occluded person ReID, we propose the Occlusion-Aware Mask Network (OAMN). In particular, we incorporate an attention-guided mask module, which requires guidance from labeled occlusion data. To this end, we propose a novel occlusion augmentation scheme that produces diverse and precisely labeled occlusion for any holistic dataset. The proposed scheme suits real-world scenarios better than existing schemes, which consider only limited types of occlusions. We also offer a novel occlusion unification scheme to tackle ambiguity information at the test phase. The above three components enable existing attention mechanisms to precisely capture body parts regardless of the occlusion. Comprehensive experiments on a variety of person ReID benchmarks demonstrate the superiority of OAMN over state-of-the-arts.
Peixian Chen, Pingyang Dai, Jianzhuang Liu, Qixiang Ye, Mingliang Xu 0001, Qi'an Chen, Rongrong Ji
ICCV3
2018 Cross-Modality Person Re-Identification with Generative Adversarial Training
abstract
Person re-identification (Re-ID) is an important task in video surveillance which automatically searches and identifies people across different cameras. Despite the extensive Re-ID progress in RGB cameras, few works have studied the Re-ID between infrared and RGB images, which is essentially a cross-modality problem and widely encountered in real-world scenarios. The key challenge lies in two folds, i.e., the lack of discriminative information to re-identify the same person between RGB and infrared modalities, and the difficulty to learn a robust metric towards such a large-scale cross-modality retrieval. In this paper, we tackle the above two challenges by proposing a novel cross-modality generative adversarial network (termed cmGAN). To handle the issue of insufficient discriminative information, we leverage the cutting-edge generative adversarial training to design our own discriminator to learn discriminative feature representation from different modalities. To handle the issue of large-scale cross-modality metric learning, we integrates both identification loss and cross-modality triplet loss, which minimize inter-class ambiguity while maximizing cross-modality similarity among instances. The entire cmGAN can be trained in an end-to-end manner by using standard deep neural network framework. We have quantized the performance of our work in the newly-released SYSU RGB-IR Re-ID benchmark, and have reported superior performance, i.e., Cumulative Match Characteristic curve (CMC) and Mean Average Precision (MAP), over the state-of-the-art works [Wu et al., 2017], respectively.
Pingyang Dai, Rongrong Ji, Qiong Wu 0012, Yuyu Huang
IJCAI1
2014 Online co-training ranking SVM for visual tracking
abstract
Online learned tracking is widely used to handle the appearance changes of object because of its adaptive ability. Learning to rank technique has attracted much attention recently in visual tracking. But the tracking method with online learning to rank suffers from the error accumulation problem during the self-training process. To solve this problem, we propose an online learning to rank algorithm in the co-training framework for robust visual tracking. A co-training algorithm combined with ranking SVM collects features and unlabeled data for training. Two ranking SVMs are built with different types of features accordingly and dynamically fused into a semi-supervised learning process. This semi-supervised learning approach is updated online to resist the occlusion and adapt to the changes of object's appearance. Many experiments on challenging sequences have shown that the proposed algorithm is more effective than the state-of-the-art methods.
Pingyang Dai, Yi Xie 0004, Cuihua Li
ICASSP1
2013 Robust visual tracking via part-based sparsity model
abstract
The sparse representation has been widely used in many areas including visual tracking. The part-based representation performs outstandingly by using non-holistic templates to against occlusion. This paper combined them and proposed a robust object tracking method using part-based sparsity model for tracking an object in a video sequence. In the proposed model, one object is represented by image patches. The candidates of these patches are sparsely represented in the space which is spanned by the patch templates and trivial templates. The part-based method takes the spatial information of each patch into consideration, where the vote maps of multiple patches are used. Furthermore, the update scheme keeps the representative templates of each part dynamically. Therefore, trackers can effectively deal with the changes of appearances and heavy occlusion. On various public benchmark videos, the abundant results of experiments demonstrate that the proposed tracking method outperforms many existing state-of-the-arts algorithms.
Pingyang Dai, Yanlong Luo, Weisheng Liu, Cuihua Li, Yi Xie 0004
ICASSP1
2013 Visual Tracking Based on Compressive Sensing MCMC Sampling
abstract
Real time visual tracking is a challenge problem in computer vision. In this paper, we propose a real-time tracking method based on compressive sensing Markov Chain Monte Carlo (MCMC) sampling. To extract the features of objects, non-adaptive random projections are employed in the object appearance model which adopts a very sparse random measurement matrix using compress sensing. These projection preserve the structure of objects in the image feature space. A Bayesian classifier is learnt from the object appearance model and the scores of this classifier are integrated into Markov Chain Monte Carlo acceptance mechanism. Furthermore, a two-stage tracking scheme is used to alleviate the drift problem. The experimental results demonstrate that the proposed method is real time and outperforms some start-of-the-art algorithms on public benchmark sequences in terms of accuracy and robustness.
Pingyang Dai, Yanlong Luo, Cuihua Li, Yi Xie 0004
SMC2
2009 Object Tracking Based on the Combination of Learning and Cascade Particle Filter
abstract
The problem of object tracking in dense clutter is a challenge in computer vision. This paper proposes a method for tracking object robustly by combining the online selection of discriminative color features and the offline selection of discriminative Haar features. Furthermore, the cascade particle filter which has four stages of importance sampling is used to fuse two kinds of features efficiently. When the illumination changes dramatically, the Haar features selected offline play a major role. When the object is occluded, or its rotation angle is very large, the color features selected online play a major role. The experimental results show that the proposed method performs well under the conditions of illumination change, occlusion, object scale change and abrupt motion of object or camera.
Hanjie Gong, Cuihua Li, Pingyang Dai, Yi Xie 0004
SMC3