De Cheng

dblp:154/1991 · DBLP profile ↗
← Back
99ranked-venue papers
31as first author
77since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 62 · 19 first-author · 51 since 2021Artificial intelligence and machine learning · 56 · 17 first-author · 43 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Harnessing Textual Semantic Priors for Knowledge Transfer and Refinement in CLIP-Driven Continual Learning
abstract
Continual learning (CL) aims to equip models with the ability to learn from a stream of tasks without forgetting previous knowledge. With the progress of vision-language models like Contrastive Language-Image Pre-training (CLIP), their promise for CL has attracted increasing attention due to their strong generalizability. However, the potential of rich textual semantic priors in CLIP in addressing the stability–plasticity dilemma remains underexplored. During backbone training, most approaches transfer past knowledge without considering semantic relevance, leading to interference from unrelated tasks that disrupt the balance between stability and plasticity. Besides, while text-based classifiers provide strong generalization, they suffer from limited plasticity due to the inherent modality gap in CLIP. Visual classifiers help bridge this gap, but their prototypes lack rich and precise semantics. To address these challenges, we propose Semantic-Enriched Continual Adaptation (SECA), a unified framework that harnesses the anti-forgetting and structured nature of textual priors to guide semantic-aware knowledge transfer in the backbone and reinforce the semantic structure of the visual classifier. Specifically, a Semantic-Guided Adaptive Knowledge Transfer (SG-AKT) module is proposed to assess new images' relevance to diverse historical visual knowledge via textual cues, and aggregate relevant knowledge in an instance-adaptive manner as distillation signals. Moreover, a Semantic-Enhanced Visual Prototype Refinement (SE-VPR) module is introduced to refine visual prototypes using inter-class semantic relations captured in class-wise textual embeddings. Extensive experiments on multiple benchmarks validate the effectiveness of our approach.
De Cheng, Di Xu 0010, Huaijie Wang, Nannan Wang 0001
AAAI2
2026 Better Matching, Less Forgetting: A Quality-Guided Matcher for Transformer-based Incremental Object Detection
abstract
Incremental Object Detection (IOD) aims to continuously learn new object classes without forgetting previously learned ones. A persistent challenge is catastrophic forgetting, primarily attributed to background shift in conventional detectors. While pseudo-labeling mitigates this in dense detectors, we identify a novel, distinct source of forgetting specific to DETR-like architectures: background foregrounding. This arises from the exhaustiveness constraint of the Hungarian matcher, which forcibly assigns every ground truth target to one prediction, even when predictions primarily cover background regions (i.e., low IoU). This erroneous supervision compels the model to misclassify background features as specific foreground classes, disrupting learned representations and accelerating forgetting. To address this, we propose a Quality-guided Min-Cost Max-Flow (Q-MCMF) matcher. To avoid forced assignments, Q-MCMF builds a flow graph and prunes implausible matches based on geometric quality. It then optimizes for the final matching that minimizes cost and maximizes valid assignments. This strategy eliminates harmful supervision from background foregrounding while maximizing foreground learning signals. Extensive experiments on the COCO dataset under various incremental settings demonstrate that our method consistently outperforms existing state-of-the-art approaches.
Qirui Wu, Shizhou Zhang, De Cheng, Yinghui Xing, Lingyan Ran, Dahu Shi, Peng Wang 0015
AAAI3
2026 A Multi-Granularity Scene-Aware Graph Convolution Method for Weakly Supervised Person Search
De Cheng, Haichun Tai, Nannan Wang 0001, Xiangqian Zhao, Jie Li 0001, Xinbo Gao 0001
Int. J. Comput. Vis.1
2026 EKPC: Elastic Knowledge Preservation and Compensation for Class-Incremental Learning
Huaijie Wang, De Cheng, Yan Li 0125, Jie Li 0001, Nannan Wang 0001, Xinbo Gao 0001
Int. J. Comput. Vis.2
2026 Isolating Interference Factors for Robust Cloth-Changing Person Re-Identification
abstract
Cloth-Changing Person Re-Identification (CC-ReID) aims to recognize individuals across camera views despite clothing variations, a crucial task for surveillance and security systems. Existing methods typically frame it as a cross-modal alignment problem but often overlook explicit modeling of interference factors such as clothing, viewpoints, and pedestrian actions. This oversight can distort their impact, compromising the extraction of robust identity features. To address these challenges, we propose a novel framework that systematically disentangles interference factors from identity features while ensuring the robustness and discriminative power of identity representations. Our approach consists of two key components. First, a dual-stream identity feature learning framework leverages a raw image stream and a cloth-isolated stream, to extract identity representations independent of clothing textures. An adaptive cloth-irrelevant contrastive objective is introduced to mitigate identity feature variations caused by clothing differences. Second, we propose a Text-Driven Conditional Generative Adversarial Interference Disentanglement Network (T-CGAIDN), to further suppress interference factors beyond clothing textures, such as finer clothing patterns, viewpoint, background, and lighting conditions. This network incorporates a multi-granularity interference recognition branch to learn interference-related features, a conditional adversarial module for bidirectional transformation between identity and interference feature spaces, and an interference decoupling objective to eliminate interference dependencies in identity learning. Extensive experiments on public benchmarks demonstrate that our method significantly outperforms state-of-the-art approaches, highlighting its effectiveness in CC-ReID.
De Cheng, Chaowei Fang, Shizhou Zhang, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2026 Prompt Disentanglement via Language Guidance and Representation Alignment for Domain Generalization
abstract
Domain Generalization (DG) seeks to develop models that perform well on unseen target domains by learning domain-invariant representations. Recent advances in pre-trained Visual Foundation Models (VFMs), such as CLIP, have shown strong potential for enhancing DG through prompt tuning. However, existing VFM-based prompt tuning methods often focus on task-specific adaptation rather than disentangling domain-invariant features, leaving cross-domain generalization insufficiently explored. In this paper, we address this challenge by fully leveraging the controllable and flexible language prompt in VFMs. Observing that the text modality is inherently rich in semantics and easier to disentangle, we propose a novel framework termed Prompt Disentanglement via Language Guidance and Representation Alignment (PADG). PADG first employs a large language model (LLM) to disentangle textual prompts into domain-invariant and domain-specific components, which then guide the learning of domain-invariant visual representations. To complement the limitations of text-only guidance, we further introduce the Worst Explicit Representation Alignment (WERA) module, which enhances visual invariance by simulating bounded domain shifts through learnable stylization prompts and aligning representations between original and perturbed samples. Extensive experiments on mainstream DG benchmarks, including PACS, VLCS, OfficeHome, DomainNet, and TerraInc, demonstrate that PADG consistently outperforms existing state-of-the-art methods, validating its effectiveness in robust domain-invariant representation learning.
De Cheng, Xinyang Jiang, Dongsheng Li 0002, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2026 VPT-NSP2++: Importance-Aware Visual Prompt Tuning in Null Space for Continual Learning
abstract
Continual learning (CL) enables AI models to adapt to evolving environments while mitigating catastrophic forgetting, which is a critical capability for dynamic real-world applications. With the growing popularity of pre-trained Vision Transformer (ViT) models and visual prompt tuning (VPT) technique in CL, this work explores a CL method on top of the ViT-based foundation model, through VPT mechanism with theoretical guarantees. Inspired by the orthogonal projection method, we aim to leverage this approach for VPT to enhance CL performance, particularly in long-term scenarios. However, since the orthogonal projection is originally designed for linear operations in CNNs, applying it to ViTs poses challenges induced by the non-linear self-attention mechanism and the distribution drift within LayerNorm. To address these issues, we deduced two orthogonality conditions to achieve the prompt gradient orthogonal projection, which provide a theoretical guarantee of maintaining stability. Considering the strict orthogonal constraints can diminish model capacity and reduce plasticity, we further propose an importance-aware orthogonal regularization framework. By applying varying degrees of orthogonal constraints to different parameters based on their importance to old and new tasks, the framework adaptively enhances model capacity and thereby promotes long-sequence CL while improving the stability-plasticity trade-off. To implement the proposed approach, a null-space-based approximation solution is employed to efficiently achieve the prompt gradient orthogonal projection. Extensive experiments on various class-incremental learning benchmarks demonstrate that our method achieves state-of-the-art performance across diverse CL scenarios.
Shizhou Zhang, Yue Lu 0008, De Cheng, Yinghui Xing, Nannan Wang 0001, Peng Wang 0015, Yanning Zhang 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2026 High-frequency structure transformer for magnetic resonance image super-resolution
Chaowei Fang, Bolin Fu, De Cheng, Lechao Cheng, Dingwen Zhang
Pattern Recognit.3
2026 Improving face forgery detection via hierarchical mixture of experts and fine-grained visual-text alignment
Chaowei Fang, Bolin Fu, De Cheng
Pattern Recognit.4
2026 Learning task-shared and specific knowledge via mixture-of-experts in generative model for continual learning
Weinan Zhao, Yanling Ji, Yan Li 0125, De Cheng, Junwei Han 0001, Dingwen Zhang
Pattern Recognit.4
2026 Semantic-Interactive Clustering Optimization With SAM for Weakly Supervised Person Search
abstract
Weakly-supervised person search presents significant challenges when relying solely on bounding-box annotations, particularly due to inter-class confusion from clothing similarity and intra-class variations caused by illumination changes, which severely degrade cross-view matching accuracy. Existing clustering-based methods, constrained by their heavy dependence on color features, frequently produce unreliable pseudo-labels that ultimately limit model performance. To overcome these limitations, we present Segment Anything Model-based Semantic-Interactive Clustering Optimization (SAM-SICO), a novel framework that integrates the Segment Anything Model’s semantic segmentation capability with adaptive clustering optimization for weakly-supervised person search. Our framework harnesses the representational power of the Segment Anything Model (SAM) to enable detector-free semantic feature learning while significantly improving clustering precision. The proposed solution makes three key advances: the Semantic Contour Embedding (SCE) module leverages SAM’s zero-shot segmentation capability to produce highly accurate human body masks; the Relation-driven Semantic Feature Interaction (RSFI) mechanism effectively mitigates clothing-color bias through innovative dynamic affinity matrix construction across multiscale semantic masks and visual features; and the Adaptive Clustering Optimization (ACO) algorithm introduces parameter adaptation to optimize intra-class compactness and inter-class separation metrics. Experimental results show that our method outperforms existing state-of-the-art approaches on the PRW and CUHK-SYSU datasets. The source code is available at https://github.com//HawlsonZ/SAM-SICO.
Xi Yang 0011, Hexun Zhou, De Cheng, Menghui Tian, Nannan Wang 0001
IEEE Trans. Circuits Syst. Video Technol.3
2026 Learning Prompt Adapters for Forgetting-Free Continual Image Super-Resolution
abstract
Continual image super-resolution (CISR) aims to efficiently adapt a pre-trained model to a variety of tasks while retaining knowledge from previously learned tasks, minimizing the need for intensive independent training. The primary challenges include catastrophic forgetting due to varying data distributions and degradation types, along with the necessity for high adaptability. While prompt-based continual learning has proven effective in image classification, its direct application to super-resolution (SR) often fails to meet the demands for detailed pixel-level restoration and domain discrimination in low-level characteristics. To address these challenges, we propose Learning Prompt Adapters (LPA), which dynamically generates pixel-wise prompts through a combination of multi-granularity prompt bases and identities. By adaptively integrating these prompts into the Transformer architecture, we effectively improve the model's performance on fine-grained details in super-resolution tasks, as well as enhancing the model's adaptability to new tasks and preserving knowledge from previous ones. Through organizing the low-rank prompt bases with specific identities, we set up an effective solution to managing cross-task differences and enhancing prompt richness. Extensive experiments on benchmarks comprising the NYU, RealSR, DIV2K, REDS, and MANGA109 datasets with diverse degradation types demonstrate that LPA significantly outperforms existing continual learning methods. Codes of this paper are available at: https://github.com/dummerchen/LPA.
Chaowei Fang, Bolin Fu, De Cheng, Chengpei Tang, Guanbin Li
IEEE Trans. Image Process.3
2026 Hierarchical Identity Learning for Unsupervised Visible-Infrared Person Re-Identification
abstract
Unsupervised visible-infrared person re-identification (USVI-ReID) aims to learn modality-invariant image features from unlabeled cross-modal person datasets by reducing the modality gap while minimizing reliance on costly manual annotations. Existing methods typically address USVI-ReID using cluster-based contrastive learning, which represents a person by a single cluster center. However, they primarily focus on the commonality of images within each cluster while neglecting the finer-grained differences among them. To address the limitation, we propose a Hierarchical Identity Learning (HIL) framework. Since each cluster may contain several smaller sub-clusters that reflect fine-grained variations among images, we generate multiple memories for each existing coarse-grained cluster via a secondary clustering. Additionally, we propose Multi-Center Contrastive Learning (MCCL) to refine representations for enhancing intra-modal clustering and minimizing cross-modal discrepancies. To further improve cross-modal matching quality, we design a Bidirectional Reverse Selection Transmission (BRST) mechanism, which establishes reliable cross-modal correspondences by performing bidirectional matching of pseudo-labels. Extensive experiments conducted on the SYSU-MM01 and RegDB datasets demonstrate that the proposed method outperforms existing approaches. The source code is available at: https://github.com/haonanshi0125/HIL.
De Cheng, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Image Process.3
2026 Multi-Level Collaborative Distillation Meets Global Workspace Model: A Unified Framework for OCIL
abstract
Online Class-Incremental Learning (OCIL) enables models to learn continuously from non-i.i.d. data streams. Since samples of the data streams can be seen only once, it is more suitable for real-world scenarios compared to offline learning. However, this constraint intensifies the challenge for OCIL in maintaining an appropriate balance between stability and plasticity. Moreover, under stricter memory buffer constraints in real world, current replay-based methods are less effective. While ensemble methods improve plasticity, they often struggle with stability. Inspired by the Global Workspace Theory (GWT), we propose a novel approach that enhances ensemble learning through a Global Workspace Model (GWM)-a shared, implicit memory that guides the learning of multiple student models. The GWM is formed by fusing the parameters of all students within each training batch, capturing the historical learning trajectory and serving as a dynamic anchor for knowledge consolidation. Like the broadcasting mechanism of GWT, the GWM is redistributed periodically to students, stabilizing learning and promoting cross-task consistency. In addition, we introduce a multi-level collaborative distillation mechanism. It enforces peer-to-peer consistency among students and preserves historical knowledge by aligning each student with the GWM. As a result, student models remain adaptable to new tasks while maintaining previously learned knowledge, striking a better balance between stability and plasticity. Extensive experiments on three standard OCIL benchmarks show that our method delivers significant performance improvement for several OCIL models across various memory budgets. The code is available at https://github.com/susususushi/GWM.
Shibin Su, Guoqiang Liang 0001, De Cheng, Shizhou Zhang, Lingyan Ran
IEEE Trans. Image Process.3
2026 ActPrompt: In-Domain Feature Adaptation via Action Cues for Video Temporal Grounding
abstract
Video temporal grounding, including moment retrieval and highlight detection, is an emerging topic aiming to identify specific clips within videos. In addition to pre-trained video models, contemporary methods utilize pre-trained vision-language models (VLMs) to capture detailed characteristics of diverse scenes and objects from video frames. However, as pre-trained on images, directly using pre-extracted VLM features neglects the domain gap between the pre-trained and temporal grounding datasets, thus inducing domain shifts due to the data-level distribution disparity. As a result, VLMs may struggle to distinguish action-sensitive patterns from static objects, making it necessary to adapt them to specific data domains for effective feature representation over temporal grounding. In this work, we address two primary challenges to achieve this goal. Specifically, to mitigate high adaptation costs, we propose an efficient preliminary in-domain fine-tuning paradigm for feature adaptation before standard downstream training, where downstream-adaptive features are learned through several well-designed pretext tasks that ensure improved performance. Furthermore, to integrate action-sensitive information into VLMs, we introduce Action-Cue-Injected Temporal Prompt Learning (ActPrompt), which injects action cues into the image encoder of VLMs to discover action-sensitive visual patterns better. This is followed by context-aware temporal prompt learning, which considers both action cues and temporal context to enhance the ability to recognize patterns associated with actions for downstream tasks. Extensive experiments demonstrate that ActPrompt is an off-the-shelf training framework that can be applied effectively to various SOTA methods, resulting in notable improvements.
Xinyang Jiang, De Cheng, Dongsheng Li 0002, Cairong Zhao
IEEE Trans. Image Process.3
2026 Overcoming Dual Incremental Challenges in Continual Person Search via Adapter and Prototype
abstract
The advancement of continual person search techniques has seen significant progress in recent years due to its practical applications in the real world. However, continual learning for person search presents significant challenges as it combines both person detection and re-identification (Re-ID) tasks, resulting in issues of domain and class incremental learning. To address these challenges, we propose a novel framework that uses an adapter-based Swin Transformer backbone, and incorporates two key components: Domain Aware Adapter (DAA) blocks and Virtual Prototype Replay-Online Instance Matching (VPR-OIM). Specifically, to solve the domain incremental problem in object detection, we introduce parallel DAA blocks to handle multiple domains, while a Domain Prototype Router (DPR) mechanism is used to dynamically route the feature to the domain-specific adapter. Additionally, for class incremental Re-ID, we extend the OIM loss with virtual prototype replay, which generates Gaussian distribution-based virtual features derived from historical prototypes, effectively enabling the model to preserve knowledge of previous identities while accommodating new identity categories. Overall, our proposed DAA and VPR-OIM simultaneously address the dual incremental challenges of continual person search. Experimental results demonstrate that our method significantly improves both person detection and Re-ID performance in continual learning settings, achieving state-of-the-art (SOTA) performance.
Xi Yang 0011, Hexun Zhou, De Cheng, Nannan Wang 0001
IEEE Trans. Image Process.3
2026 Dual-Domain Adaptation Networks for Realistic Image Super-Resolution
abstract
Realistic image super-resolution (SR) focuses on transforming real-world low-resolution (LR) images into high-resolution (HR) ones, handling more complex degradation patterns than synthetic SR tasks. This is critical for applications like surveillance, medical imaging, and consumer electronics. However, current methods struggle with limited real-world LR-HR data, impacting the learning of basic image features. Pre-trained SR models from large-scale synthetic datasets offer valuable prior knowledge, which can improve generalization, speed up training, and reduce the need for extensive real-world data in realistic SR tasks. In this paper, we introduce a novel approach,Dual-domain Adaptation Networks, which is able to efficiently adapt pre-trained image SR models from simulated to real-world datasets. To achieve this target, we first set up a spatial-domain adaptation strategy through selectively updating parameters of pre-trained models and employing the low-rank adaptation technique to adjust frozen parameters. Recognizing that image super-resolution involves recovering high-frequency components, we further integrate a frequency domain adaptation branch into the adapted model, which combines the spectral data of the input and the spatial-domain backbone's intermediate features to infer HR frequency maps, enhancing the SR result. Experimental evaluations on public realistic image SR benchmarks, including RealSR, D2CRealSR, and DRealSR, demonstrate the superiority of our proposed method over existing state-of-the-art models.
Chaowei Fang, Bolin Fu, De Cheng, Lechao Cheng, Guanbin Li
IEEE Trans. Multim.3
2025 Dual Information Purification for Lightweight SAR Object Detection
abstract
Synthetic aperture radar (SAR) object detection requires accurate identification and localization of targets at various scales within SAR images. However, background clutter and speckle noise can obscure key features and mislead the knowledge distillation process. To address these challenges, we introduce the Dual Information Purification Knowledge Distillation (DIPKD) method, which improves the performance of the student model through three key strategies: denoising, enrichment, and decoupling. First, our Selective Noise Suppression (SNS) technique reduces speckle noise in global features by minimizing misleading information from the teacher model. Second, the Knowledge Level Decoupling (KLD) module separates features into target and non-target knowledge, balancing feature mapping and reducing background noise to enhance the extraction of critical information for the student model. Finally, the Reverse Information Transfer (RIT) module refines intermediate features in the student model, compensating for the loss of detailed local information. Experimental results demonstrate that DIPKD significantly outperforms existing distillation techniques in SAR object detection, achieving 60.2% and 51.4% mAP scores on the SSDD and HRSID datasets, respectively. Additionally, the student model shows performance improvements of 1.3% and 2.9% over the teacher model, highlighting the effectiveness of the information purification approach.
Xi Yang 0011, Songsong Duan, De Cheng
AAAI4
2025 Training Consistent Mixture-of-Experts-Based Prompt Generator for Continual Learning
abstract
Visual prompt tuning-based continual learning (CL) methods have shown promising performance in exemplar-free scenarios, where their key component can be viewed as a prompt generator. Existing approaches generally rely on freezing old prompts, slow updating and task discrimination for prompt generators to preserve stability and minimize forgetting. In contrast, we introduce a novel approach that trains a consistent prompt generator to ensure stability during CL. Consistency means that for any instance from an old task, its corresponding instance-ware prompt generated by the prompt generator remains consistent even as the generator continually updates in a new task. This ensures that the representation of a specific instance remains stable across tasks and thereby prevents forgetting. We employ a mixture of experts (MoE) as the prompt generator, which contains a router and multiple experts. By deriving conditions sufficient to achieve the consistency for the MoE prompt generator, we demonstrate that: during training in a new task, if the router and experts update in the directions orthogonal to the subspaces spanned by old input features and gating vectors, respectively, the consistency can be theoretically guaranteed. To implement this orthogonality, we project parameter gradients to those orthogonal directions using the orthogonal projection matrices computed via the null space method. Extensive experiments on four class-incremental learning benchmarks validate the effectiveness and superiority of our approach.
Yue Lu 0008, Shizhou Zhang, De Cheng, Guoqiang Liang 0001, Yinghui Xing, Nannan Wang 0001, Yanning Zhang 0001
AAAI3
2025 Adversarial Domain Prompt Tuning and Generation for Single Domain Generalization
abstract
Single domain generalization (SDG) aims to learn a robust model, which could perform well on many unseen domains while there is only one single domain available for training. One of the promising directions for achieving single-domain generalization is to generate out-of-domain (OOD) training data through data augmentation or image generation. Given the rapid advancements in AI-generated content (AIGC), this paper is the first to propose leveraging powerful pre-trained text-to-image (T2I) foundation models to create the training data. However, manually designing textual prompts to generate images for all possible domains is often impractical, and some domain characteristics may be too abstract to describe with words. To address these challenges, we propose a novel Progressive Adversarial Prompt Tuning (PAPT) framework for pre-trained diffusion models. Instead of relying on static textual domains, our approach learns two sets of abstract prompts as conditions for the diffusion model: one that captures domain-invariant category information and another that models domain-specific styles. This adversarial learning mechanism enables the T2I model to generate images in various domain styles while preserving key categorical features. Extensive experiments demonstrate the effectiveness of the proposed method, achieving superior performances to state-of-the-art single-domain generalization approaches.
De Cheng, Xinyang Jiang, Nannan Wang 0001, Dongsheng Li 0002, Xinbo Gao 0001
CVPR2
2025 Top-Push Polynomial Ranking Embedded Dictionary Learning for Enhanced Re-Id
Ying Chen 0041, De Cheng, Zhihui Li 0001, Andy Song
ICAART (3)2
2025 Gradient Decomposition and Alignment for Incremental Object Detection
Wenlong Luo, Shizhou Zhang, De Cheng, Yinghui Xing, Guoqiang Liang 0001, Peng Wang 0015, Yanning Zhang 0001
ICCV3
2025 Dual Domain Control via Active Learning for Remote Sensing Domain Incremental Object Detection
De Cheng, Xi Yang 0011, Nannan Wang 0001
ICCV2
2025 Demystifying Catastrophic Forgetting in Two-Stage Incremental Object Detector
abstract
Catastrophic forgetting is a critical chanllenge for incremental object detection (IOD). Most existing methods treat the detector monolithically, relying on instance replay or knowledge distillation without analyzing component-specific forgetting. Through dissection of Faster R-CNN, we reveal a key insight: Catastrophic forgetting is predominantly localized to the RoI Head classifier, while regressors retain robustness across incremental stages. This finding challenges conventional assumptions, motivating us to develop a framework termed NSGP-RePRE. Regional Prototype Replay (RePRE) mitigates classifier forgetting via replay of two types of prototypes: coarse prototypes represent class-wise semantic centers of RoI features, while fine-grained prototypes model intra-class variations. Null Space Gradient Projection (NSGP) is further introduced to eliminate prototype-feature misalignment by updating the feature extractor in directions orthogonal to subspace of old inputs via gradient projection, aligning RePRE with incremental learning dynamics. Our simple yet effective design allows NSGP-RePRE to achieve state-of-the-art performance on the Pascal VOC and MS COCO datasets under various settings. Our work not only advances IOD methodology but also provide pivotal insights for catastrophic forgetting mitigation in IOD. Code will be available soon.
Qirui Wu, Shizhou Zhang, De Cheng, Yinghui Xing, Di Xu 0010, Peng Wang 0015, Yanning Zhang 0001
ICML3
2025 Screening, Rectifying, and Re-Screening: A Unified Framework for Tuning Vision-Language Models with Noisy Labels
abstract
Pre-trained vision-language models have shown remarkable potential for downstream tasks. However, their fine-tuning under noisy labels remains an open problem due to challenges like self-confirmation bias and the limitations of conventional small-loss criteria. In this paper, we propose a unified framework to address these issues, consisting of three key steps: Screening, Rectifying, and Re-Screening. First, a dual-level semantic matching mechanism is introduced to categorize samples into clean, ambiguous, and noisy samples by leveraging both macro-level and micro-level textual prompts. Second, we design tailored pseudo-labeling strategies to rectify noisy and ambiguous labels, enabling their effective incorporation into the training process. Finally, a re-screening step, utilizing cross-validation with an auxiliary vision-language model, mitigates self-confirmation bias and enhances the robustness of the framework. Extensive experiments across ten datasets demonstrate that the proposed method significantly outperforms existing approaches for tuning vision-language pre-trained models with noisy labels.
Chaowei Fang, Hangfei Ma, De Cheng, Yue Zhang 0025, Guanbin Li
IJCAI4
2025 Boosting Multi-Modal Alignment: Geometric Feature Separation for Class Incremental Learning
abstract
Class Incremental Learning (CIL) aims to continually learn new classes from a stream of data without forgetting previously learned ones. Recent approaches have leveraged pre-trained models (PTMs) to improve performance, especially vision-language models, which offer better generalization than models trained solely on visual data. Many of these methods rely on simple language templates to generate class representations, which then serve as classifiers. However, due to differences between the pre-training data and downstream tasks, these textual features can become too similar for certain classes, leading to prediction errors. To address this issue, we propose a method that optimizes the geometric structure of both visual and textual features across different classes. Inspired by neural collapse theory, we introduce a multi-modal alignment strategy: for each class, a reference vector is chosen from a simplex Equiangular Tight Frame, and both the visual and textual features of the class are aligned with this vector. To better capture intra-class variations, we also construct multiple visual prototypes for each class. A multi-prototype supervised contrastive loss is then employed to pull an image feature toward the closest matching prototype of its true class and push it away from prototypes of other classes. We evaluate our approach on five widely used CIL benchmarks. The results show that our method achieves state-of-the-art performance, demonstrating its effectiveness in addressing the challenges of class incremental learning. Our code is available at https://github.com/qcNPU/NCSCMP.
Guoqiang Liang 0001, De Cheng, Shizhou Zhang, Yanning Zhang 0001
ACM Multimedia3
2025 Amplitude-aware Domain Style Replay for Lifelong Person Re-identification
abstract
Lifelong Person Re-identification (LReID) focuses on continuously adapting to new domains over time while preserving knowledge from previously seen domains, particularly under the domain incremental learning setting. The major challenge of LReID is catastrophic forgetting, typically caused by large domain shifts during training. To address this, we propose a novel Amplitude-aware Domain Style Replay (ADSR) framework, which introduces a Fourier-based Style Transfer (FST) mechanism to generate synthetic data that reflects the style of previously encountered domains. These proxy images help retain prior knowledge without the need to store actual past data. Our method transfers stylistic information-mainly encoded in the amplitude spectrum-from old domains to new ones, creating old-stylized images that preserve the content of new domain data while adopting the visual style of earlier domains. To further boost generalization, we design a Self-Stylization Normalization (SSN) module that adapts the current domain's style distribution, making the model more robust to stylistic variations. Additionally, we introduce a Multi-Granularity Transfer (MGT) module that uses K-Means clustering to extract multiple representative style features from each domain, enabling compact yet comprehensive storage and replay of domain-specific information. Extensive experiments on multiple LReID benchmarks show that ADSR achieves superior performance over existing approaches, effectively reducing forgetting and improving cross-domain generalization. Our code is available at https://github.com/cclong8/MM2025-ADSR.
De Cheng, Shizhou Zhang, Yinghui Xing, Di Xu 0010, Yanning Zhang 0001
ACM Multimedia2
2025 Semantic-Aligned Learning with Collaborative Refinement for Unsupervised VI-ReID
De Cheng, Nannan Wang 0001, Dingwen Zhang, Xinbo Gao 0001
Int. J. Comput. Vis.1
2025 Exploring Homogeneous and Heterogeneous Consistent Label Associations for Unsupervised Visible-Infrared Person ReID
De Cheng, Nannan Wang 0001, Xinbo Gao 0001
Int. J. Comput. Vis.2
2025 Image dehazing via self-supervised depth guidance
Yudong Liang, Shaoji Li, De Cheng, Wenjian Wang 0001, Deyu Li 0001, Jiye Liang
Pattern Recognit.3
2025 Achieving Plasticity-Stability Trade-Off in Continual Learning Through Adaptive Orthogonal Projection
abstract
Catastrophic forgetting is the crucial challenge for continual learning. One of the state-of-the-art approaches is the orthogonal projection, which aims to learn each task by updating model parameters in the direction orthogonal to the subspace spanned by the previous task input. Although such strict orthogonal weight constraints ensure no interference with tasks that have been learned to achieve model stability, they greatly sacrifice model plasticity. In this paper, we propose an adaptive balanced orthogonal projection (AdaBOP) method, to search for the optimal network parameter updating direction to address the plasticity-stability dilemma in continual learning. The proposed AdaBOP method can adaptively adjust its tendency towards plasticity-stability trade-off based on the layer-wise feature space correlations of the model between old and new tasks. To further improve the training efficiency, we also implement the AdaBOP method in the uncentered covariance matrix space of the previous tasks, and finally achieve a better stability-plasticity trade-off in continual learning efficiently. Experimental results greatly demonstrate the effectiveness of the proposed method, which achieves superior performances to state-of-the-art continual learning approaches. The code is available athttps://github.com/hyscn/AdaBOP.
De Cheng, Yusong Hu, Nannan Wang 0001, Dingwen Zhang, Xinbo Gao 0001
IEEE Trans. Circuits Syst. Video Technol.1
2025 Progressive Feature-Attribute Matching via Bi-Directional Generation for Transductive Zero-Shot Learning
abstract
Transductive zero-shot learning (TZSL) has been proposed to address the domain shift problem by leveraging additional unlabeled unseen data to enhance the generalization ability from seen classes to unseen target classes. Existing TZSL methods primarily focus on mitigating the distribution bias problem by incorporating these unlabeled samples into the generative models. Although these methods have achieved great success, they do not fully exploit the potential of these unlabeled target data. In this paper, we propose a bidirectional weakly guided conditional generative modeling approach, which utilizes the attribute regressor and the visual generator to synthesize paired training data of unseen classes for each other, thus converting unlabeled target data into matched feature-attribute pairs. Additionally, on top of the generative modeling, we also propose to progressively estimate the associations between visual features and attributes among the unlabeled target data through a semi-supervised pseudo-labeling approach, so as to further facilitate the generative model and enhance the learning of target distributions. Extensive experimental results on four benchmark datasets demonstrate the effectiveness of the proposed method, achieving superior performances to state-of-the-art methods. Our source code is released in https://github.com/LevisWei/semi-zero-master.
De Cheng, Chaowei Fang, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Circuits Syst. Video Technol.1
2025 CatVersion: Concatenating Embeddings for Diffusion-Based Text-to-Image Personalization
abstract
We propose CatVersion, an inversion-based method that learns the personalized concept through a handful of examples. Subsequently, users can utilize text prompts to generate images that embody the personalized concept, thereby achieving text-to-image personalization. In contrast to existing approaches that emphasize word embedding learning or parameter fine-tuning for the diffusion model, which potentially causes concept dilution or overfitting, our method concatenates embeddings on the feature-dense space of the text encoder in the diffusion model to learn the gap between the personalized concept and its base class, aiming to maximize the preservation of prior knowledge in diffusion models while restoring the personalized concepts. To this end, we first dissect the text encoder’s integration in the image generation process to identify the feature-dense space of the encoder. Afterward, we concatenate embeddings on the Keys and Values in this space to learn the gap between the personalized concept and its base class. In this way, the concatenated embeddings ultimately manifest as a residual on the original attention output. To more accurately and unbiasedly quantify the results of personalized image generation, we improve the CLIP image alignment score based on masks. Qualitatively and quantitatively, CatVersion helps to restore personalization concepts more faithfully and enables more robust editing.
Mingrui Zhu, Shiyin Dong, De Cheng, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Circuits Syst. Video Technol.4
2025 Escaping Modal Interactions: An Efficient DESANet for Multi-Modal Object Re-Identification
abstract
Multi-modal object Re-ID aims to leverage the complementary information provided by multiple modalities to overcome challenging conditions and achieve high-quality object matching. However, existing multi-modal methods typically rely on various modality interaction modules for information fusion, which can reduce the efficiency of real-time monitoring systems. Additionally, practical challenges such as low-quality multi-modal data or missing modalities further complicate the application of object Re-ID. To address these issues, we propose the Complementary Data Enhancement and Modal-Aware Soft Alignment Network (DESANet), which is designed to be independent of interactive networks and adaptable to scenarios with missing modalities. This approach ensures a simple-yet-effective, and efficient multi-modal object Re-ID. DESANet consists of three key components: Firstly, the Dual-Color Space Data Enhancement (DCDE) module, which enhances multi-modal data by performing patch rotation in the RGB space and improving image quality in the HSV space. Secondly, the Salient Feature ReConstruction (SFRC) module, which addresses the issue of missing modalities by reconstructing features from one modality using the other two. Thirdly, the Modal-Aware Soft Alignment (MASA) module, which integrates multi-source data to avoid the blind fusion of features and prevents the propagation of noise from reconstructed modalities. Our approach achieves state-of-the-art performances on both person and vehicle datasets. Source code is available at https://github.com/DWJ11/DESANet.
Wenjiao Dong, Xi Yang 0011, De Cheng, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Image Process.3
2025 Enhancing Feature Learning With Hard Samples in Mutual Learning for Online Class Incremental Learning
abstract
Online Class-Incremental Learning (OCIL) aims to solve the problem of incrementally learning new classes from a non-i.i.d. and single-pass data stream. Compared to the offline setting, OCIL is much closer to a live learning experience requiring higher model update frequency at less computational budget. Due to its one-epoch training constraint, the model is likely to learn non-essential features and encounter the under-fitting issue, which severely affects the model's stability. In this paper, we investigate how to use hard samples to improve data variability, eventually enhancing feature learning and addressing the under-fitting problem. Specifically, by introducing a scoring function assessing the sample value, we build an OCIL formulation that simultaneously generates high-value samples and optimizes the OCIL model, improving generalization ability within the constraint of single-epoch training. Empirically, we found that strong data augmentation is a simple but effective way to generate a higher proportion of high-score samples. To make the most of these augmented samples, we design an OCIL model based on mutual learning with two networks of identical structures. Moreover, a collaborative learning mechanism is developed by aligning the features and class probabilities from the two networks to promote their interaction. Extensive experiments on three widely used datasets for OCIL have demonstrated the effectiveness of our method, obtaining superior performance to state-of-the-art methods. The code is available at https://github.com/susususushi/SDA-MCL.
Guoqiang Liang 0001, Shibin Su, De Cheng, Shizhou Zhang, Peng Wang 0015, Yanning Zhang 0001
IEEE Trans. Image Process.3
2025 FA-Net: A Feature Alignment Network for Video-Based Visible-Infrared Person Re-Identification
abstract
Video-based visible-infrared person re-identification (VVI-ReID) aims to match target pedestrians between visible and infrared videos, which is significantly applied in 24-hour surveillance systems. The key of VVI-ReID is to learn modality invariant and spatio-temporal invariant sequence-level representation to solve the challenges such as modality differences, spatio-temporal misalignment, and domain shift noise. However, existing methods predominantly emphasize on reducing modality discrepancy while relatively neglect temporal misalignment and domain shift noise reduction. To this end, this paper proposes a VVI-ReID framework called Feature Alignment Network (FA-Net) from the perspective of feature alignment, aiming to mitigate temporal misalignment. FA-Net comprises two main alignment modules: Spatial-Temporal Alignment Module (STAM) and Modality Distribution Constraint (MDC). STAM integrates global and local features to ensure individuals' spatial representation alignment. Additionally, STAM also establishes temporal relationships by exploring inter-frame features to address cross-frame person feature matching. Furthermore, we introduce the Modality Distribution Constraint (MDC), which utilizes a symmetric distribution loss to align the distributions of features from different modalities. Besides, the SAM Guidance Augmentation (SAM-GA) strategy is designed to transform the image space of RGB and IR frames to provide more informative and less noisy frame information. Extensive experimental results demonstrate the effectiveness of the proposed method, surpassing existing state-of-the-art methods. Our code will be available at: https://github.com/code/FANet.
Xi Yang 0011, Wenjiao Dong, De Cheng, Nannan Wang 0001
IEEE Trans. Image Process.4
2025 Prompt-Based Modality Alignment for Effective Multi-Modal Object Re-Identification
abstract
A critical challenge for multi-modal Object Re-Identification (ReID) is the effective aggregation of complementary information to mitigate illumination issues. State-of-the-art methods typically employ complex and highly-coupled architectures, which unavoidably result in heavy computational costs. Moreover, the significant distribution gap among different image spectra hinders the joint representation of multi-modal features. In this paper, we propose a framework named as PromptMA to establish effective communication channels between different modality paths, thereby aggregating modal complementary information and bridging the distribution gap. Specifically, we inject a series of learnable multi-modal prompts into the Image Encoder and introduce a prompt exchange mechanism to enable the prompts to alternately interact with different modal token embeddings, thus capturing and distributing multi-modal features effectively. Building on top of the multi-modal prompts, we further propose Prompt-based Token Selection (PBTS) and Prompt-based Modality Fusion (PBMF) modules to achieve effective multi-modal feature fusion while minimizing background interference. Additionally, due to the flexibility of our prompt exchange mechanism, our method is well-suited to handle scenarios with missing modalities. Extensive evaluations are conducted on four widely used benchmark datasets and the experimental results demonstrate that our method achieves state-of-the-art performances, surpassing the current benchmarks by over 15% on the challenging MSVR310 dataset and by 6% on the RGBNT201. The code is available at https://github.com/FHR-L/PromptMA.
Shizhou Zhang, Wenlong Luo, De Cheng, Yinghui Xing, Guoqiang Liang 0001, Peng Wang 0015, Yanning Zhang 0001
IEEE Trans. Image Process.3
2025 ETC: Temporal Boundary Expand Then Clarify for Weakly Supervised Video Grounding With Multimodal Large Language Model
abstract
Early weakly supervised video grounding (WSVG) methods often struggle with incomplete boundary detection due to the absence of temporal boundary annotations. To bridge the gap between video-level and boundary-level annotations, explicit supervision methods (i.e., generating pseudo-temporal boundaries for training) have achieved great success. However, data augmentation in these methods might disrupt critical temporal information, yielding poor pseudo-temporal boundaries. In this paper, we propose a new perspective that maintains the integrity of the original temporal content while introducing more valuable information for expanding the incomplete boundaries. To this end, we proposeETC(ExpandthenClarify), first using the additional information to expand the initial incomplete pseudo-temporal boundaries, and subsequently refining these expanded ones to achieve precise boundaries. Motivated by video continuity, i.e., visual similarity across adjacent frames, we use powerful multi-modal large language models (MLLMs) to annotate each frame within the initial pseudo-temporal boundaries, yielding more comprehensive descriptions for expanded boundaries. To further clarify the noise in expanded boundaries, we combine mutual learning with a tailored proposal-level contrastive objective to use a learnable approach to harmonize a balance between incomplete yet clean (initial) and comprehensive yet noisy (expanded) boundaries for more precise ones. Experiments demonstrate the superiority of our method on two challenging WSVG datasets.
Guozhang Li, Xinpeng Ding, De Cheng, Jie Li 0001, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Multim.3
2025 Progressive Prompt-Driven Low-Light Image Enhancement With Frequency Aware Learning
abstract
Low-light Image Enhancement (LLIE) aims to rectify inadequate illumination conditions and achieve superior visual quality in images, which plays a pivotal role in the domain of low-level computer vision. Due to poor illumination in images, many high-frequency details are obscured, which leads to an uneven distribution of low- and high-frequency information. However, most existing LLIE methods do not pay special attention to the restoration of high-frequency detail information and some challenging-to-recover areas in images. To address this issue, we propose a novel progressive prompt-driven LLIE framework with frequency aware learning, through a two-stage coarse-to-fine learning mechanism. Specifically, the proposed method fully utilizes both the specially designed brightness-aware prompt and detail-aware prompt on the prior trained model, to achieve an excellent enhanced image that exhibits more natural brightness and richer detail information. Furthermore, the proposed frequency aware learning objective can adaptively adjust the contribution of individual pixels for image reconstruction based on the statistics of high- and low-frequency features, which enables the network to focus on learning intricate details and other challenging areas in low-light images. Extensive experimental results demonstrate the effectiveness of the proposed method, achieving superior performances to state-of-the-art methods on representative real-world and synthetic datasets. Our source code is available athttps://github.com/MSL502/PPFAL.
De Cheng, Yan Li 0125, Nannan Wang 0001, Dingwen Zhang, Xinbo Gao 0001, Jiande Sun 0001
IEEE Trans. Multim.2
2025 Capsule Networks With Residual Pose Routing
abstract
Capsule networks (CapsNets) have been known difficult to develop a deeper architecture, which is desirable for high performance in the deep learning era, due to the complex capsule routing algorithms. In this article, we present a simple yet effective capsule routing algorithm, which is presented by a residual pose routing. Specifically, the higher-layer capsule pose is achieved by an identity mapping on the adjacently lower-layer capsule pose. Such simple residual pose routing has two advantages: 1) reducing the routing computation complexity and 2) avoiding gradient vanishing due to its residual learning framework. On top of that, we explicitly reformulate the capsule layers by building a residual pose block. Stacking multiple such blocks results in a deep residual CapsNets (ResCaps) with a ResNet-like architecture. Results on MNIST, AffNIST, SmallNORB, and CIFAR-10/100 show the effectiveness of ResCaps for image classification. Furthermore, we successfully extend our residual pose routing to large-scale real-world applications, including 3-D object reconstruction and classification, and 2-D saliency dense prediction. The source code has been released on https://github.com/liuyi1989/ResCaps.
Yi Liu 0038, De Cheng, Dingwen Zhang, Shoukun Xu, Jungong Han
IEEE Trans. Neural Networks Learn. Syst.2
2025 TIENet: A Tri-Interaction Enhancement Network for Multimodal Person Reidentification
abstract
Multimodal person reidentification (ReID), which aims to learn modality-complementary information by utilizing multimodal images simultaneously for person retrieval, is crucial for achieving all-time and all-weather monitoring. Existing methods try to address this issue through modality fusion to absorb complementary information. However, most of these methods are limited to the spatial domain only and usually overlook the intra-/intermodal interactions during feature fusion, resulting in insufficient learning of modality-specific and complementary information. To address these issues, we propose a tri-interaction enhancement network (TIENet), which contains three modules: spatial-frequency interaction (SFI), intermodal mask interaction (IMMI), and intramodal feature fusion (IMFF). Specifically, the SFI boosts the modality-specific representation by integrating the amplitude-guided attention mechanism into the phase space, combined with spatial-domain convolution to achieve fine-grained information learning. Meanwhile, the IMMI enhances the richness of the feature descriptors by embedding the intermodal relationships to preserve complementary information. Finally, the IMFF module considers the structure of the human body and integrates intramodal contextual information. Extensive experimental results demonstrate the effectiveness of our method, achieving superior performances on RGBNT201 and MARKET1501_RGBNT datasets.
Xi Yang 0011, Wenjiao Dong, De Cheng, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Neural Networks Learn. Syst.3
2024 Learning Hierarchical Prompt with Structured Linguistic Knowledge for Vision-Language Models
abstract
Prompt learning has become a prevalent strategy for adapting vision-language foundation models to downstream tasks. As large language models (LLMs) have emerged, recent studies have explored the use of category-related descriptions as input to enhance prompt effectiveness. Nevertheless, conventional descriptions fall short of structured information that effectively represents the interconnections among entities or attributes linked to a particular category. To address this limitation and prioritize harnessing structured knowledge, this paper advocates for leveraging LLMs to build a graph for each description to model the entities and attributes describing the category, as well as their correlations. Preexisting prompt tuning methods exhibit inadequacies in managing this structured knowledge. Consequently, we propose a novel approach called Hierarchical Prompt Tuning (HPT), which enables simultaneous modeling of both structured and conventional linguistic knowledge. Specifically, we introduce a relationship-guided attention module to capture pair-wise associations among entities and attributes for low-level prompt learning. In addition, by incorporating high-level and global-level prompts modeling overall semantics, the proposed hierarchical structure forges cross-level interlinks and empowers the model to handle more complex and long-term relationships. Extensive experiments demonstrate that our HPT shows strong effectiveness and generalizes much better than existing SOTA methods. Our code is available at https://github.com/Vill-Lab/2024-AAAI-HPT.
Xinyang Jiang, De Cheng, Dongsheng Li 0002, Cairong Zhao
AAAI3
2024 Disentangled Prompt Representation for Domain Generalization
abstract
Domain Generalization (DG) aims to develop a versatile model capable of performing well on unseen target domains. Recent advancements in pre-trained Visual Foundation Models (VFMs), such as CLIP, show significant potential in enhancing the generalization abilities of deep models. Although there is a growing focus on VFM-based domain prompt tuning for DG, effectively learning prompts that disentangle invariant features across all domains remains a major challenge. In this paper, we propose addressing this challenge by leveraging the controllable and flexible language prompt of the VFM. Observing that the text modality of VFMs is inherently easier to disentangle, we introduce a novel text feature guided visual prompt tuning framework. This framework first automatically disentangles the text prompt using a large language model (LLM) and then learns domain-invariant visual representation guided by the disentangled text feature. Moreover, we also devise domain-specific prototype learning to fully exploit domain-specific information to combine with the invariant feature prediction. Extensive experiments on mainstream DG datasets, namely PACS, VLCS, OfficeHome, DomainNet and TerraInc, demonstrate that the proposed method achieves superior performances to state-of-the-art DG methods.
De Cheng, Xinyang Jiang, Nannan Wang 0001, Dongsheng Li 0002, Xinbo Gao 0001
CVPR1
2024 Cross-Platform Video Person ReID: A New Benchmark Dataset and Adaptation Approach
Shizhou Zhang, Wenlong Luo, De Cheng, Qingchun Yang, Lingyan Ran, Yinghui Xing, Yanning Zhang 0001
ECCV (27)3
2024 Gradient and Brightness Guided Low-Light Enhancement with Attention-Based Self-Paced Learning
abstract
Low-light image enhancement aims to reconstruct images with insufficient illumination into visually appealing representations with natural brightness. While most existing methods tend to focus on enhancing illumination, they often overlook the restoration of finer details in the enhanced image. Moreover, these methods do not adequately address the varying degradation levels observed in different regions of the image. In this study, we present a gradient and brightness guided low-light image enhancement framework, which can simultaneously augment the detail and illumination during the enhancement process. Our approach involves extracting gradient information from gamma-corrected images, which offers a remarkable advantage in preserving edge details compared to direct extraction from degraded images. To further refine the enhancement process and adaptively adjust the difficulty of samples, thereby boosting learning efficiency, we introduce an attention-based self-paced learning strategy. This strategy assigns different gradient and brightness weights based on the degradation levels within different image regions. Extensive experiments demonstrate the superiority of our proposed method over state-of-the-art approaches. The code is available at https://github.com/MSL502/GBASPL.
Yan Li 0125, De Cheng, Dingwen Zhang, Luofeng Zhai, Jiande Sun 0001
ICASSP3
2024 Task-aware Orthogonal Sparse Network for Exploring Shared Knowledge in Continual Learning
abstract
Continual learning (CL) aims to learn from sequentially arriving tasks without catastrophic forgetting (CF). By partitioning the network into two parts based on the Lottery Ticket Hypothesis—one for holding the knowledge of the old tasks while the other for learning the knowledge of the new task—the recent progress has achieved forget-free CL. Although addressing the CF issue well, such methods would encounter serious under-fitting in long-term CL, in which the learning process will continue for a long time and the number of new tasks involved will be much higher. To solve this problem, this paper partitions the network into three parts—with a new part for exploring the knowledge sharing between the old and new tasks. With the shared knowledge, this part of network can be learnt to simultaneously consolidate the old tasks and fit to the new task. To achieve this goal, we propose a task-aware Orthogonal Sparse Network (OSN), which contains shared knowledge induced network partition and sharpness-aware orthogonal sparse network learning. The former partitions the network to select shared parameters, while the latter guides the exploration of shared knowledge through shared parameters. Qualitative and quantitative analyses, show that the proposed OSN induces minimum to no interference with past tasks, i.e., approximately no forgetting, while greatly improves the model plasticity and capacity, and finally achieves the state-of-the-art performances.
Yusong Hu, De Cheng, Dingwen Zhang, Nannan Wang 0001, Tongliang Liu, Xinbo Gao 0001
ICML2
2024 Dual-Branch Task Residual Enhancement with Parameter-Free Attention for Zero-Shot Multi-label Image Recognition
Shizhou Zhang, Kairui Dang, De Cheng, Yinghui Xing, Qirui Wu, Dexuan Kong, Yanning Zhang 0001
ICPR (22)3
2024 Multi-Granularity Graph-Convolution-Based Method for Weakly Supervised Person Search
Haichun Tai, De Cheng, Jie Li 0001, Nannan Wang 0001, Xinbo Gao 0001
IJCAI2
2024 Disentangling Identity Features from Interference Factors for Cloth-Changing Person Re-identification
abstract
Cloth-Changing Person Re-Identification (CC-ReID) aims to accurately identify a target person in the more realistic surveillance scenario where clothes of the pedestrian may change drastically, which is critical in public security systems for tracking down disguised criminal suspects. Existing methods mainly transform the CC-ReID problem into cross-modality feature alignment from the data-driven perspective, without modelling the interference factors such as clothes and camera view changes meticulously. This may lead to over-consideration or under-consideration of the influence of these factors on the extraction of robust and discriminative identity features. This paper proposes a novel algorithm for thoroughly disentangling identity features from interference factors brought by clothes and camera view changes while ensuring the robustness and discriminability. It adopts a dual-stream identity feature learning framework consisting of a raw image stream and a cloth-erasing stream, to explore discriminative and cloth-irrelevant identity feature representations. Specifically, an adaptive cloth-irrelevant contrastive objective is introduced to contrast features extracted by the two streams, aiming to suppress the fluctuation caused by clothes textures in the identity feature space. Moreover, we innovatively mitigate the influence of the interference factors through a generative adversarial interference factor decoupling network. This network is targeted at capturing identity-related information residing in the interference factors and disentangling the identity features from such information. Extensive experimental results demonstrate the effectiveness of the proposed method, achieving superior performances to state-of-the-art methods.
De Cheng, Chaowei Fang, Changzhe Jiao, Nannan Wang 0001, Xinbo Gao 0001
ACM Multimedia2
2024 Feature-Level Adversarial Attacks and Ranking Disruption for Visible-Infrared Person Re-identification
abstract
Visible-infrared person re-identification (VIReID) is widely used in fields such as video surveillance and intelligent transportation, imposing higher demands on model security. In practice, the adversarial attacks based on VIReID aim to disrupt output ranking and quantify the security risks of models. Although numerous studies have been emerged on adversarial attacks and defenses in fields such as face recognition, person re-identification, and pedestrian detection, there is currently a lack of research on the security of VIReID systems. To this end, we propose to explore the vulnerabilities of VIReID systems and prevent potential serious losses due to insecurity. Compared to research on single-modality ReID, adversarial feature alignment and modality differences need to be particularly emphasized. Thus, we advocate for feature-level adversarial attacks to disrupt the output rankings of VIReID systems. To obtain adversarial features, we introduce \textit{Universal Adversarial Perturbations} (UAP) to simulate common disturbances in real-world environments. Additionally, we employ a \textit{Frequency-Spatial Attention Module} (FSAM), integrating frequency information extraction and spatial focusing mechanisms, and further emphasize important regional features from different domains on the shared features. This ensures that adversarial features maintain consistency within the feature space. Finally, we employ an \textit{Auxiliary Quadruple Adversarial Loss} to amplify the differences between modalities, thereby improving the distinction and recognition of features between visible and infrared images, which causes the system to output incorrect rankings. Extensive experiments on two VIReID benchmarks (i.e., SYSU-MM01, RegDB) and different systems validate the effectiveness of our method.
Xi Yang 0011, De Cheng, Nannan Wang 0001, Xinbo Gao 0001
NeurIPS3
2024 Visual Prompt Tuning in Null Space for Continual Learning
abstract
Existing prompt-tuning methods have demonstrated impressive performances in continual learning (CL), by selecting and updating relevant prompts in the vision-transformer models. On the contrary, this paper aims to learn each task by tuning the prompts in the direction orthogonal to the subspace spanned by previous tasks' features, so as to ensure no interference on tasks that have been learned to overcome catastrophic forgetting in CL. However, different from the orthogonal projection in the traditional CNN architecture, the prompt gradient orthogonal projection in the ViT architecture shows completely different and greater challenges, i.e., 1) the high-order and non-linear self-attention operation; 2) the drift of prompt distribution brought by the LayerNorm in the transformer block. Theoretically, we have finally deduced two consistency conditions to achieve the prompt gradient orthogonal projection, which provide a theoretical guarantee of eliminating interference on previously learned knowledge via the self-attention mechanism in visual prompt tuning. In practice, an effective null-space-based approximation solution has been proposed to implement the prompt gradient orthogonal projection. Extensive experimental results demonstrate the effectiveness of anti-forgetting on four class-incremental benchmarks with diverse pre-trained baseline models, and our approach achieves superior performances to state-of-the-art methods. Our code is available at https://github.com/zugexiaodui/VPTinNSforCL
Yue Lu 0008, Shizhou Zhang, De Cheng, Yinghui Xing, Nannan Wang 0001, Peng Wang 0015, Yanning Zhang 0001
NeurIPS3
2024 Diffusion-based Layer-wise Semantic Reconstruction for Unsupervised Out-of-Distribution Detection
abstract
Unsupervised out-of-distribution (OOD) detection aims to identify out-of-domain data by learning only from unlabeled In-Distribution (ID) training samples, which is crucial for developing a safe real-world machine learning system. Current reconstruction-based method provides a good alternative approach, by measuring the reconstruction error between the input and its corresponding generative counterpart in the pixel/feature space. However, such generative methods face the key dilemma, $i.e.$, improving the reconstruction power of the generative model, while keeping compact representation of the ID data. To address this issue, we propose the diffusion-based layer-wise semantic reconstruction approach for unsupervised OOD detection. The innovation of our approach is that we leverage the diffusion model's intrinsic data reconstruction ability to distinguish ID samples from OOD samples in the latent feature space. Moreover, to set up a comprehensive and discriminative feature representation, we devise a multi-layer semantic feature extraction strategy. Through distorting the extracted features with Gaussian noises and applying the diffusion model for feature reconstruction, the separation of ID and OOD samples is implemented according to the reconstruction errors. Extensive experimental results on multiple benchmarks built upon various datasets demonstrate that our method achieves state-of-the-art performance in terms of detection accuracy and speed.
Ying Yang 0020, De Cheng, Chaowei Fang, Yubiao Wang, Changzhe Jiao, Lechao Cheng, Nannan Wang 0001, Xinbo Gao 0001
NeurIPS2
2024 M-RRFS: A Memory-Based Robust Region Feature Synthesizer for Zero-Shot Object Detection
Peiliang Huang, Dingwen Zhang, De Cheng, Longfei Han, Pengfei Zhu 0001, Junwei Han 0001
Int. J. Comput. Vis.3
2024 Efficient Statistical Sampling Adaptation for Exemplar-Free Class Incremental Learning
abstract
Deep learning systems typically suffer from catastrophic forgetting of old knowledge when learning from new data continually. Recently, various class incremental learning (CIL) methods have been proposed to address this issue, and some approaches achieve promising performances by relying on rehearsing the training data of previous tasks. However, storing data from previous tasks would encounter data privacy and memory issues in real-world applications. In this paper, we propose a statistical sampling adaptation method for efficient Exemplar-Free Class-Incremental Learning (EFCIL). Here, instead of preserving the images/features themselves of previous tasks/classes, we store image feature statistics from previous classes to maintain the decision boundary, which is memory-efficient and much semantic-representative. When utilizing the old-class feature statistics, we build a statistical feature adaptation network (SFAN) with a manifold consistency regularization and then train it in a transductive learning paradigm, which can map the outdated statistics onto the current feature space to facilitate a compatible and balanced classifier training subsequently. In this way, the final classifier can be jointly optimized with all the old-class features projected by SFAN and current new-class features, thus alleviating the classification bias problem in EFCIL. Experimental results greatly demonstrate the effectiveness of the proposed method, achieving superior performances than state-of-the-art approaches. Our source code is released inhttps://github.com/yxzhcv/ESSA-EFCIL.
De Cheng, Nannan Wang 0001, Guozhang Li, Dingwen Zhang, Xinbo Gao 0001
IEEE Trans. Circuits Syst. Video Technol.1
2024 Neighbor Consistency and Global-Local Interaction: A Novel Pseudo-Label Refinement Approach for Unsupervised Person Re-Identification
abstract
Unsupervised person re-identification (ReID) aims at learning discriminative identity features for person retrieval without any annotations. Recent advances accomplish this task by leveraging clustering-based pseudo labels, but these pseudo labels are inevitably noisy, which deteriorates model performance. In this paper, we propose a Neighbour Consistency guided Pseudo Label Refinement (NCPLR) framework, which can be regarded as a transductive form of label propagation under the assumption that the prediction of each example should be similar to its nearest neighbours’. Specifically, the refined label for each training instance can be obtained from the original clustering result and a weighted ensemble of its neighbours’ predictions, with weights determined according to their similarities in the feature space. Furthermore, we also explore building a unified global-local NCPLR mechanism through a global-local label interaction module to achieve mutual label refinement. Such a strategy promotes efficient complementary learning while mitigating some unreliable information, finally improving the quality of the refined pseudo labels for each global-local region. Extensive experimental results demonstrate the effectiveness of the proposed method, showing superior performance to state-of-the-art methods by a large margin. Our source code is released inhttps://github.com/haichuntai/NCPLR-ReID.
De Cheng, Haichun Tai, Nannan Wang 0001, Chaowei Fang, Xinbo Gao 0001
IEEE Trans. Inf. Forensics Secur.1
2024 Neighbor-Guided Pseudo-Label Generation and Refinement for Single-Frame Supervised Temporal Action Localization
abstract
Due to the sparse single-frame annotations, current Single-Frame Temporal Action Localization (SF-TAL) methods generally employ threshold-based pseudo-label generation strategies. However, these approaches suffer from inefficient data utilization, as only parts of unlabeled frames with confidence scores surpassing a predefined threshold are selected for training. Moreover, the variability of single-frame annotations and unreliable model predictions introduce pseudo-label noise. To address these challenges, we propose two strategies by using the relationship of the video segments with their neighbors': 1) temporal neighbor-guided soft pseudo-label generation (TNPG); and 2) semantic neighbor-guided pseudo-label refinement (SNPR). TNPG utilizes a local-global self-attention mechanism in a transformer encoder to capture temporal neighbor information while focusing on the whole video. Then the generated self-attention map is multiplied by the network predictions to propagate information between labeled and unlabeled frames, and produce soft pseudo-label for all segments. Despite this, label noise persists due to unreliable model predictions. To mitigate this, SNPR refines pseudo-labels based on the assumption that predictions should resemble their semantic nearest neighbors'. Specifically, we search for semantic nearest neighbors of each video segment by cosine similarity in the feature space. Then the refined soft pseudo-labels can be obtained by a weight combination of the original pseudo-label and the semantic nearest neighbors'. Finally, the model can be trained with the refined pseudo-labels, and the performance has been greatly improved. Comprehensive experimental results on different benchmarks show that we achieve state-of-the-art performances on THUMOS14, ActivityNet1.2, and ActivityNet1.3 datasets.
Guozhang Li, De Cheng, Nannan Wang 0001, Jie Li 0001, Xinbo Gao 0001
IEEE Trans. Image Process.2
2024 Continual All-in-One Adverse Weather Removal With Knowledge Replay on a Unified Network Structure
abstract
In real-world applications, image degeneration caused by adverse weather is always complex and changes with different weather conditions from days and seasons. Systems in real-world environments constantly encounter adverse weather conditions that are not previously observed. Therefore, it practically requires adverse weather removal models to continually learn from incrementally collected data reflecting various degeneration types. Existing adverse weather removal approaches, for either single or multiple adverse weathers, are mainly designed for a static learning paradigm, which assumes that the data of all types of degenerations to handle can be finely collected at one time before a single-phase learning process. They thus cannot directly handle the incremental learning requirements. To address this issue, we made the earliest effort to investigate the continual all-in-one adverse weather removal task, in a setting closer to real-world applications. Specifically, we develop a novel continual learning framework with effective knowledge replay (KR) on a unified network structure. Equipped with a principal component projection and an effective knowledge distillation mechanism, the proposed KR techniques are tailored for the all-in-one weather removal task. It considers the characteristics of the image restoration task with multiple degenerations in continual learning, and the knowledge for different degenerations can be shared and accumulated in the unified network structure. Extensive experimental results demonstrate the effectiveness of the proposed method to deal with this challenging task, which performs competitively to existing dedicated or joint training image restoration methods. Our code is available athttps://github.com/xiaojihh/CL_all-in-one.
De Cheng, Yanling Ji, Dong Gong, Yan Li 0125, Nannan Wang 0001, Junwei Han 0001, Dingwen Zhang
IEEE Trans. Multim.1
2024 Progressive Negative Enhancing Contrastive Learning for Image Dehazing and Beyond
abstract
Image dehazing is a pivotal preliminary step in the advancement of robust intelligent surveillance system. However, it is an extremely challenging ill-posed problem, as it faces severe information degradation when accurately restoring the clean image from its haze-polluted counterpart. This paper proposes a novel Progressive Negative Enhancing (PNE) contrastive learning mechanism to fully exploit various types of negative information, thereby facilitating the traditional positive-oriented objective function for image dehazing. The proposed method can progressively update the negative samples during model training, to steadily squeeze the restored image towards its desired clean target from various directions. Furthermore, considering the image dehazing task as a many-to-one feature mapping problem, we also make an early effort to enhance the robustness of the dehazing model under variational haze densities. Specifically, a novel density-variational dehazing network is proposed to be optimized under the consistency-regularized framework using the proposed PNE learning mechanism. The consistency regularization ensures consistent output given multi-level degraded hazy images, thereby significantly enhancing the robustness of the model in dealing with various hazy scenarios. Extensive experiments demonstrate that the proposed method exhibits superior performance over existing state-of-the-art methods. It achieves average PSNR boosts of 0.60dB, 0.28dB and 0.82dB on dehazing, deraining and desnowing tasks, respectively. The source code is available athttps://github.com/YanLi-LY/PNE-Net.
De Cheng, Yan Li 0125, Dingwen Zhang, Nannan Wang 0001, Jiande Sun 0001, Xinbo Gao 0001
IEEE Trans. Multim.1
2024 Dual Modality Prompt Tuning for Vision-Language Pre-Trained Model
abstract
With the emergence of large pretrained vison-language models such as CLIP, transferable representations can be adapted to a wide range of downstream tasks via prompt tuning. Prompt tuning probes for beneficial information for downstream tasks from the general knowledge stored in the pretrained model. A recently proposed method named Context Optimization (CoOp) introduces a set of learnable vectors as text prompts from the language side. However, tuning the text prompt alone can only adjust the synthesized “classifier”, while the computed visual features of the image encoder cannot be affected, thus leading to suboptimal solutions. In this article, we propose a novel dual-modality prompt tuning (DPT) paradigm through learning text and visual prompts simultaneously. To make the final image feature concentrate more on the target visual concept, a class-aware visual prompt tuning (CAVPT) scheme is further proposed in our DPT. In this scheme, the class-aware visual prompt is generated dynamically by performing the cross attention between text prompt features and image patch token embeddings to encode both the downstream task-related information and visual instance information. Extensive experimental results on 11 datasets demonstrate the effectiveness and generalization ability of the proposed method.
Yinghui Xing, Qirui Wu, De Cheng, Shizhou Zhang, Guoqiang Liang 0001, Peng Wang 0015, Yanning Zhang 0001
IEEE Trans. Multim.3
2024 Weakly Supervised Temporal Action Localization With Bidirectional Semantic Consistency Constraint
abstract
Weakly supervised temporal action localization (WTAL) aims to classify and localize temporal boundaries of actions for the video, given only video-level category labels in the training datasets. Due to the lack of boundary information during training, existing approaches formulate WTAL as a classification problem, i.e., generating the temporal class activation map (T-CAM) for localization. However, with only classification loss, the model would be suboptimized, i.e., the action-related scenes are enough to distinguish different class labels. Regarding other actions in the action-related scene (i.e., the scene same as positive actions) as co-scene actions, this suboptimized model would misclassify the co-scene actions as positive actions. To address this misclassification, we propose a simple yet efficient method, named bidirectional semantic consistency constraint (Bi-SCC), to discriminate the positive actions from co-scene actions. The proposed Bi-SCC first adopts a temporal context augmentation to generate an augmented video that breaks the correlation between positive actions and their co-scene actions in the inter-video. Then, a semantic consistency constraint (SCC) is used to enforce the predictions of the original video and augmented video to be consistent, hence suppressing the co-scene actions. However, we find that this augmented video would destroy the original temporal context. Simply applying the consistency constraint would affect the completeness of localized positive actions. Hence, we boost the SCC in a bidirectional way to suppress co-scene actions while ensuring the integrity of positive actions, by cross-supervising the original and augmented videos. Finally, our proposed Bi-SCC can be applied to current WTAL approaches and improve their performance. Experimental results show that our approach outperforms the state-of-the-art methods on THUMOS14 and ActivityNet. The code is available at https://github.com/lgzlIlIlI/BiSCC.
Guozhang Li, De Cheng, Xinpeng Ding, Nannan Wang 0001, Jie Li 0001, Xinbo Gao 0001
IEEE Trans. Neural Networks Learn. Syst.2
2023 Cross-Modality Person Re-identification with Memory-Based Contrastive Embedding
abstract
Visible-infrared person re-identification (VI-ReID) aims to retrieve the person images of the same identity from the RGB to infrared image space, which is very important for real-world surveillance system. In practice, VI-ReID is more challenging due to the heterogeneous modality discrepancy, which further aggravates the challenges of traditional single-modality person ReID problem, i.e., inter-class confusion and intra-class variations. In this paper, we propose an aggregated memory-based cross-modality deep metric learning framework, which benefits from the increasing number of learned modality-aware and modality-agnostic centroid proxies for cluster contrast and mutual information learning. Furthermore, to suppress the modality discrepancy, the proposed cross-modality alignment objective simultaneously utilizes both historical and up-to-date learned cluster proxies for enhanced cross-modality association. Such training mechanism helps to obtain hard positive references through increased diversity of learned cluster proxies, and finally achieves stronger ``pulling close'' effect between cross-modality image features. Extensive experiment results demonstrate the effectiveness of the proposed method, surpassing state-of-the-art works significantly by a large margin on the commonly used VI-ReID datasets.
De Cheng, Nannan Wang 0001, Zhen Wang 0037, Xiaoyu Wang 0002, Xinbo Gao 0001
AAAI1
2023 Boosting Weakly-Supervised Temporal Action Localization with Text Information
abstract
Due to the lack of temporal annotation, current Weakly-supervised Temporal Action Localization (WTAL) methods are generally stuck into over-complete or incomplete localization. In this paper, we aim to leverage the text information to boost WTAL from two aspects, i.e., (a) the discriminative objective to enlarge the inter-class difference, thus reducing the over-complete; (b) the generative objective to enhance the intra-class integrity, thus finding more complete temporal boundaries. For the discriminative objective, we propose a Text-Segment Mining (TSM) mechanism, which constructs a text description based on the action class label, and regards the text as the query to mine all class-related segments. Without the temporal annotation of actions, TSM compares the text query with the entire videos across the dataset to mine the best matching segments while ignoring irrelevant ones. Due to the shared sub-actions in different categories of videos, merely applying TSM is too strict to neglect the semantic-related segments, which results in incomplete localization. We further introduce a generative objective named Video-text Language Completion (VLC), which focuses on all semantic-related segments from videos to complete the text sentence. We achieve the state-of-the-art performance on THUMOS14 and ActivityNetl.3. Surprisingly, we also find our proposed method can be seamlessly applied to existing methods, and improve their performances with a clear margin. The code is available at https://github.com/lgzlIlIlI/Boosting-WTAL.
Guozhang Li, De Cheng, Xinpeng Ding, Nannan Wang 0001, Xiaoyu Wang 0002, Xinbo Gao 0001
CVPR2
2023 Unsupervised Visible-Infrared Person ReID by Collaborative Learning with Neighbor-Guided Label Refinement
abstract
Unsupervised learning visible-infrared person re-identification (USL-VI-ReID) aims at learning modality-invariant features from unlabeled cross-modality dataset, which is crucial for practical applications in video surveillance systems. The key to essentially address the USL-VI-ReID task is to solve the cross-modality data association problem for further heterogeneous joint learning. To address this issue, we propose a Dual Optimal Transport Label Assignment (DOTLA) framework to simultaneously assign the generated labels from one modality to its counterpart modality. The proposed DOTLA mechanism formulates a mutual reinforcement and efficient solution to cross-modality data association, which could effectively reduce the side-effects of some insufficient and noisy label associations. Besides, we further propose a cross-modality neighbor consistency guided label refinement and regularization module, to eliminate the negative effects brought by the inaccurate supervised signals, under the assumption that the prediction or label distribution of each example should be similar to its nearest neighbors'. Extensive experimental results on the public SYSU-MM01 and RegDB datasets demonstrate the effectiveness of the proposed method, surpassing existing state-of-the-art approach by a large margin of 7.76% mAP on average, which even surpasses some supervised VI-ReID methods.
De Cheng, Xiaojian Huang, Nannan Wang 0001, Zhihui Li 0001, Xinbo Gao 0001
ACM Multimedia1
2023 Efficient Bilateral Cross-Modality Cluster Matching for Unsupervised Visible-Infrared Person ReID
abstract
Unsupervised visible-infrared person re-identification (USL-VI-ReID) aims to match pedestrian images of the same identity from different modalities without annotations. Existing works mainly focus on alleviating the modality gap by aligning instance-level features of the unlabeled samples. However, the relationships between cross-modality clusters are not well explored. To this end, we propose a novel bilateral cluster matching-based learning framework to reduce the modality gap by matching cross-modality clusters. Specifically, we design a Many-to-many Bilateral Cross-Modality Cluster Matching (MBCCM) algorithm through optimizing the maximum matching problem in a bipartite graph. Then, the matched pairwise clusters utilize shared visible and infrared pseudo-labels during the model training. Under such a supervisory signal, a Modality-Specific and Modality-Agnostic (MSMA) contrastive learning framework is proposed to align features jointly at a cluster-level. Meanwhile, the cross-modality Consistency Constraint (CC) is proposed to explicitly reduce the large modality discrepancy. Extensive experiments on the public SYSU-MM01 and RegDB datasets demonstrate the effectiveness of the proposed method, surpassing state-of-the-art approaches by a large margin of 8.76% mAP on average.
De Cheng, Nannan Wang 0001, Shizhou Zhang, Zhen Wang 0037, Xinbo Gao 0001
ACM Multimedia1
2023 Ground-to-Aerial Person Search: Benchmark Dataset and Approach
abstract
In this work, we construct a large-scale dataset for Ground-to-Aerial Person Search, named G2APS, which contains 31,770 images of 260,559 annotated bounding boxes for 2,644 identities appearing in both of the UAVs and ground surveillance cameras. To our knowledge, this is the first dataset for cross-platform intelligent surveillance applications, where the UAVs could work as a powerful complement for the ground surveillance cameras. To more realistically simulate the actual cross-platform Ground-to-Aerial surveillance scenarios, the surveillance cameras are fixed about 2 meters above the ground, while the UAVs capture videos of persons at different location, with a variety of view-angles, flight attitudes and flight modes. Therefore, the dataset has the following unique characteristics: 1) drastic view-angle changes between query and gallery person images from cross-platform cameras; 2) diverse resolutions, poses and views of the person images under 9 rich real-world scenarios. On basis of the G2APS benchmark dataset, we demonstrate detailed analysis about current two-step and end-to-end person search methods, and further propose a simple yet effective knowledge distillation scheme on the head of the ReID network, which achieves state-of-the-art performances on both of the G2APS and the previous two public person search datasets, i.e., PRW and CUHK-SYSU. The dataset and source code available on https://github.com/yqc123456/HKD_for_person_search.
Shizhou Zhang, Qingchun Yang, De Cheng, Yinghui Xing, Guoqiang Liang 0001, Peng Wang 0015, Yanning Zhang 0001
ACM Multimedia3
2023 Hybrid routing transformer for zero-shot learning
De Cheng, Gerong Wang, Bo Wang 0011, Qiang Zhang 0020, Jungong Han, Dingwen Zhang
Pattern Recognit.1
2023 Discriminative and Robust Attribute Alignment for Zero-Shot Learning
abstract
Zero-shot learning (ZSL) aims to learn models that can recognize images of semantically related unseen categories, through transferring attribute-based knowledge learned from training data of seen classes to unseen testing data. As visual attributes play a vital role in ZSL, recent embedding-based methods usually focus on learning a compatibility function between the visual representation and the class semantic attributes. While in this work, in addition to simply learning the region embedding of different semantic attributes to maintain the generalization capability of the learned model, we further consider to improve the discrimination power of the learned visual features themselves by contrastive embedding. It exploits both the class-wise and instance-wise supervision for GZSL, under the attribute guided weakly supervised representation learning framework. To further improve the robustness of the ZSL model, we also propose to train the model under the consistency regularization constraint, through taking full advantages of self-supervised signals of the image under various perturbed augmentation situations, which could make the model robust to some occluded or un-related attribute regions. Extensive experimental results demonstrate the effectiveness of the proposed ZSL method, achieving superior performances to state-of-the-art methods on three widely-used benchmark datasets, namely CUB, SUN, and AWA2. Our source code is released athttps://github.com/KORIYN/CC-ZSL.
De Cheng, Gerong Wang, Nannan Wang 0001, Dingwen Zhang, Qiang Zhang 0020, Xinbo Gao 0001
IEEE Trans. Circuits Syst. Video Technol.1
2022 Instance-Dependent Label-Noise Learning with Manifold-Regularized Transition Matrix Estimation
abstract
In label-noise learning, estimating the transition matrix has attracted more and more attention as the matrix plays an important role in building statistically consistent classifiers. However, it is very challenging to estimate the transition matrix T(x), where x denotes the instance, because it is unidentifiable under the instance-dependent noise (IDN). To address this problem, we have noticed that, there are psychological and physiological evidences showing that we humans are more likely to annotate instances of similar appearances to the same classes, and thus poor-quality or ambiguous instances of similar appearances are easier to be mislabeled to the correlated or same noisy classes. Therefore, we propose assumption on the geometry of T(x) that “the closer two instances are, the more similar their corresponding transition matrices should be”. More specifically, we formulate above assumption into the manifold embedding, to effectively reduce the degree of freedom of T(x) and make it stably estimable in practice. The proposed manifold-regularized technique works by directly reducing the estimation error without hurting the approximation error about the estimation problem of T(x). Experimental evaluations on four synthetic and two real-world datasets demonstrate that our method is superior to state-of-the-art approaches for label-noise learning under the challenging IDN.
De Cheng, Tongliang Liu, Yixiong Ning, Nannan Wang 0001, Bo Han 0003, Gang Niu 0001, Xinbo Gao 0001, Masashi Sugiyama
CVPR1
2022 Robust Region Feature Synthesizer for Zero-Shot Object Detection
abstract
Zero-shot object detection aims at incorporating class semantic vectors to realize the detection of (both seen and) unseen classes given an unconstrained test image. In this study, we reveal the core challenges in this research area: how to synthesize robust region features (for unseen objects) that are as intra-class diverse and inter-class separable as the real samples, so that strong unseen object detectors can be trained upon them. To address these challenges, we build a novel zero-shot object detection framework that contains an Intra-class Semantic Diverging component and an Inter-class Structure Preserving component. The former is used to realize the one-to-more mapping to obtain diverse visual features from each class semantic vector, preventing miss-classifying the real unseen objects as image backgrounds. While the latter is used to avoid the synthesized features too scattered to mix up the inter-class and foreground-background relationship. To demonstrate the effectiveness of the proposed approach, comprehensive experiments on PASCAL VOC, COCO, and DIOR datasets are conducted. Notably, our approach achieves the new state-of-the-art performance on PASCAL VOC and COCO and it is the first study to carry out zero-shot object detection in remote sensing imagery.
Peiliang Huang, Junwei Han 0001, De Cheng, Dingwen Zhang
CVPR3
2022 Robust Single Image Dehazing Based on Consistent and Contrast-Assisted Reconstruction
abstract
Single image dehazing as a fundamental low-level vision task, is essential for the development of robust intelligent surveillance system. In this paper, we make an early effort to consider dehazing robustness under variational haze density, which is a realistic while under-studied problem in the research filed of singe image dehazing. To properly address this problem, we propose a novel density-variational learning framework to improve the robustness of the image dehzing model assisted by a variety of negative hazy images, to better deal with various complex hazy scenarios. Specifically, the dehazing network is optimized under the consistency-regularized framework with the proposed Contrast-Assisted Reconstruction Loss (CARL). The CARL can fully exploit the negative information to facilitate the traditional positive-orient dehazing objective function, by squeezing the dehazed image to its clean target from different directions. Meanwhile, the consistency regularization keeps consistent outputs given multi-level hazy images, thus improving the model robustness. Extensive experimental results on two synthetic and three real-world datasets demonstrate that our method significantly surpasses the state-of-the-art approaches.
De Cheng, Yan Li 0125, Dingwen Zhang, Nannan Wang 0001, Xinbo Gao 0001, Jiande Sun 0001
IJCAI1
2022 Class-Dependent Label-Noise Learning with Cycle-Consistency Regularization
abstract
In label-noise learning, estimating the transition matrix plays an important role in building statistically consistent classifier. Current state-of-the-art consistent estimator for the transition matrix has been developed under the newly proposed sufficiently scattered assumption, through incorporating the minimum volume constraint of the transition matrix T into label-noise learning. To compute the volume of T, it heavily relies on the estimated noisy class posterior. However, the estimation error of the noisy class posterior could usually be large as deep learning methods tend to easily overfit the noisy labels. Then, directly minimizing the volume of such obtained T could lead the transition matrix to be poorly estimated. Therefore, how to reduce the side-effects of the inaccurate noisy class posterior has become the bottleneck of such method. In this paper, we creatively propose to estimate the transition matrix under the forward-backward cycle-consistency regularization, of which we have greatly reduced the dependency of estimating the transition matrix T on the noisy class posterior. We show that the cycle-consistency regularization helps to minimize the volume of the transition matrix T indirectly without exploiting the estimated noisy class posterior, which could further encourage the estimated transition matrix T to converge to its optimal solution. Extensive experimental results consistently justify the effectiveness of the proposed method, on reducing the estimation error of the transition matrix and greatly boosting the classification performance.
De Cheng, Yixiong Ning, Nannan Wang 0001, Xinbo Gao 0001, Bo Han 0003, Tongliang Liu
NeurIPS1
2022 Single image dehazing with an independent Detail-Recovery Network
Yan Li 0125, De Cheng, Dingwen Zhang, Nannan Wang 0001, Xinbo Gao 0001, Jiande Sun 0001
Knowl. Based Syst.2
2022 Learning Deep Resonant Prior for Hyperspectral Image Super-Resolution
abstract
Hyperspectral image super-resolution (HSISR) task has been widely studied, and significant progress has been made by leveraging the deep convolution neural network (CNN) techniques. Nevertheless, the scarcity of training images hinders the research progress of HSISR task. Moreover, the differences in imaging conditions and the number of spectral bands among different datasets, make it very difficult to construct a unified deep neural network. In this paper, we first present a non-training based HSISR method based on deep prior knowledge, which captures the image prior to restore the high resolution image by using the intrinsic characteristics of CNN. Then, we append a special network input processing module onto the HSI super-resolution network to automatically adjust the structure of the input so that the choice of network structure is no longer limited, while the network design focuses on exploiting the spatial information of hyperspectral images and the correlation between spectral bands, making the method more suitable for HSISR tasks and greatly extending its applications. Extensive experiment results on the hyperspectral image datasets illustrate the effectiveness of the proposed method, and we have got comparable results with the state-of-the-art methods while requiring no training samples.
Zhaori Gong, Nannan Wang 0001, De Cheng, Jingwei Xin, Xi Yang 0011, Xinbo Gao 0001
IEEE Trans. Geosci. Remote. Sens.3
2022 Hybrid Dynamic Contrast and Probability Distillation for Unsupervised Person Re-Id
abstract
Unsupervised person re-identification (Re-Id) has attracted increasing attention due to its practical application in the read-world video surveillance system. The traditional unsupervised Re-Id are mostly based on the method alternating between clustering and fine-tuning with the classification or metric learning objectives on the grouped clusters. However, since person Re-Id is an open-set problem, the clustering based methods often leave out lots of outlier instances or group the instances into the wrong clusters, thus they can not make full use of the training samples as a whole. To solve these problems, we present the hybrid dynamic cluster contrast and probability distillation algorithm. It formulates the unsupervised Re-Id problem into an unified local-to-global dynamic contrastive learning and self-supervised probability distillation framework. Specifically, the proposed method can make the best of the self-supervised signals of all the clustered and un-clustered instances, from both the instances' self-contrastive level and the probability distillation respectives, in the memory-based non-parametric manner. Besides, the proposed hybrid local-to-global contrastive learning can take full advantage of the informative and valuable training examples for effective and robust training. Extensive experiment results show that the proposed method achieves superior performances to state-of-the-art methods, under both the purely unsupervised and unsupervised domain adaptation experiment settings. Our source code is released in https://github.com/zjy2050/HDCRL-ReID.
De Cheng, Jingyu Zhou, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Image Process.1
2021 Lossless Compression for Video Streams with Frequency Prediction and Macro Block Merging
abstract
Cloud service has been emerging as a promising alternative to handle massive volumes of video sequences triggered by increasing demands on video service, especially surveillance and entertainment. Lossless compression of encoded video bitstreams can further eliminate the redundancies without altering the contents and facilitate the efficiency of cloud storage. In this paper, we propose a novel lossless compression scheme to further compress the video bitstreams generated by the state-of-the-art hybrid coding frameworks like H.264/AVC and HEVC. Different from transcoding, the proposed scheme develops frequency prediction and Macro Block (MB) merging to eliminate the redundancies remained in the intra-and inter-predicted frames with a strict guarantee of video fidelity. To our best knowledge, this paper is the first attempt to realize lossless compression of video bitstreams generated by advanced coding standards H.264/AVC and HEVC. Experimental results demonstrate that the proposed scheme can achieve compression gains of 17.38% and 10.85% on standard test sequences and surveillance videos, respectively.
Jixiang Luo, Wenrui Dai, De Cheng, Junni Zou, Hongkai Xiong
DCC4
2021 Support-Set Based Cross-Supervision for Video Grounding
abstract
Current approaches for video grounding propose kinds of complex architectures to capture the video-text relations, and have achieved impressive improvements. However, it is hard to learn the complicated multi-modal relations by only architecture designing in fact. In this paper, we introduce a novel Support-set Based Cross-Supervision (Sscs) module which can improve existing methods during training phase without extra inference cost. The proposed Sscs module contains two main components, i.e., discriminative contrastive objective and generative caption objective. The contrastive objective aims to learn effective representations by contrastive learning, while the caption objective can train a powerful video encoder supervised by texts. Due to the co-existence of some visual entities in both ground-truth and background intervals, i.e. mutual exclusion, naively contrastive learning is unsuitable to video grounding. We address the problem by boosting the cross-supervision with the support-set concept, which collects visual information from the whole video and eliminates the mutual exclusion of entities. Combined with the original objectives, Sscs can enhance the abilities of multi-modal relation modeling for existing approaches. We extensively evaluate Sscs on three challenging datasets, and show that our method can improve current state-of-the-art methods by large margins, especially 6.35% in terms of [email protected] on Charades-STA.
Xinpeng Ding, Nannan Wang 0001, Shiwei Zhang 0001, De Cheng, Xiaomeng Li 0001, Ziyuan Huang 0003, Mingqian Tang, Xinbo Gao 0001
ICCV4
2021 Disentangling Deep Network for Reconstructing 3D Object Shapes from Single 2D Images
Yang Yang 0009, Junwei Han 0001, Dingwen Zhang, De Cheng
PRCV (2)4
2020 Noise-to-Compression Variational Autoencoder for Efficient End-to-End Optimized Image Coding
abstract
Generative model has emerged as a disruptive alternative for lossy compression of natural images, but suffers from the low-fidelity reconstruction. In this paper, we propose a noise-to-compression variational antoencoder (NC-VAE) to achieve efficient rate-distortion optimization (RDO) for end-to-end optimized image compression with a guarantee of fidelity. The proposed NC-VAE improves rate-distortion performance by adaptively adjusting the distribution of latent variables with trainable noise perturbation. Consequently, high-efficiency RDO is developed based on the distribution of latent variables for simplified decoder. Furthermore, robust end-to-end learning is developed over the corrupted inputs to suppress the deformation and color drift in standard VAE based generative models. Experimental results show that NC-VAE outperforms the state-of-the-art lossy image coders and recent end-to-end optimized compression methods in low bit-rate region, i.e., below 0.2 bits per pixel (bpp).
Jixiang Luo, Wenrui Dai, Yuhui Xu 0002, De Cheng, Hongkai Xiong
DCC5
2020 A model of co-saliency based audio attention
XiaoMing Zhao, De Cheng
Multim. Tools Appl.3
2020 Object detection with class aware region proposal network and focused attention objective
Yihong Gong, Weiwei Shi 0003, De Cheng
Pattern Recognit. Lett.4
2020 Fusion of Multiple Person Re-id Methods With Model and Data-Aware Abilities
abstract
Person re-identification (person re-id) has attracted rapidly increasing attention in computer vision and pattern recognition research community in recent years. With the goal of providing match ranking results between each query person image and the gallery ones, the person re-id technique has been widely explored and a large number of person re-id methods have been developed. As these algorithms leverage different kinds of prior assumptions, image features, distance matching functions, et al., each of them has its own strengths and weaknesses. Inspired by these facts, this paper proposes a novel person re-id method based on the idea of inferring superior fusion results from a variety of previous base person re-id algorithms using different methodologies or features. To achieve this goal, we propose a novel framework which mainly consists of two steps: 1) a number of existing person re-id methods are implemented, and the ranking results are obtained in the test datasets. and 2) the robust fusion strategy is applied to obtain better re-ranked matching results by simultaneously considering the recognition abilities of various base re-id methods and the difficulties of different gallery person images to be correctly recognized under the generative model of labels, abilities, and difficulties framework. Comprehensive experiments show the effectiveness of our proposed method, and we have received state-of-the-art results on recent popular person re-id datasets.
De Cheng, Zhihui Li 0001, Yihong Gong, Dingwen Zhang
IEEE Trans. Cybern.1
2019 Consistency-Preserving deep hashing for fast person re-identification
Diangang Li, Yihong Gong, De Cheng, Weiwei Shi 0003, Xinyuan Chang
Pattern Recognit.3
2019 Fine-Grained Image Classification Using Modified DCNNs Trained by Cascaded Softmax and Generalized Large-Margin Losses
abstract
We develop a fine-grained image classifier using a general deep convolutional neural network (DCNN). We improve the fine-grained image classification accuracy of a DCNN model from the following two aspects. First, to better model the h -level hierarchical label structure of the fine-grained image classes contained in the given training data set, we introduce h fully connected (fc) layers to replace the top fc layer of a given DCNN model and train them with the cascaded softmax loss. Second, we propose a novel loss function, namely, generalized large-margin (GLM) loss, to make the given DCNN model explicitly explore the hierarchical label structure and the similarity regularities of the fine-grained image classes. The GLM loss explicitly not only reduces between-class similarity and within-class variance of the learned features by DCNN models but also makes the subclasses belonging to the same coarse class be more similar to each other than those belonging to different coarse classes in the feature space. Moreover, the proposed fine-grained image classification framework is independent and can be applied to any DCNN structures. Comprehensive experimental evaluations of several general DCNN models (AlexNet, GoogLeNet, and VGG) using three benchmark data sets (Stanford car, fine-grained visual classification-aircraft, and CUB-200-2011) for the fine-grained image classification task demonstrate the effectiveness of our method.
Weiwei Shi 0003, Yihong Gong, De Cheng, Nanning Zheng 0001
IEEE Trans. Neural Networks Learn. Syst.4
2018 Pedestrian search in surveillance videos by learning discriminative deep features
Shizhou Zhang, De Cheng, Yihong Gong, Dahu Shi, Xi Qiu, Yong Xia 0001, Yanning Zhang 0001
Neurocomputing2
2018 Person re-identification by the asymmetric triplet and identification loss function
De Cheng, Yihong Gong, Weiwei Shi 0003, Shizhou Zhang
Multim. Tools Appl.1
2018 Correction to: Person re-identification by the symmetric triplet and identification loss function
De Cheng, Yihong Gong, Weiwei Shi 0003, Shizhou Zhang
Multim. Tools Appl.1
2018 View-invariant gait recognition based on kinect skeleton feature
Jiande Sun 0001, Jing Li 0046, Wenbo Wan, De Cheng, Huaxiang Zhang 0001
Multim. Tools Appl.5
2018 Deep feature learning via structured graph Laplacian embedding for person re-identification
De Cheng, Yihong Gong, Xiaojun Chang, Weiwei Shi 0003, Alex Hauptmann 0001, Nanning Zheng 0001
Pattern Recognit.1
2018 Entropy and orthogonality based deep discriminative feature learning for object recognition
Weiwei Shi 0003, Yihong Gong, De Cheng, Nanning Zheng 0001
Pattern Recognit.3
2018 Two-Stream Multirate Recurrent Neural Network for Video-Based Pedestrian Reidentification
abstract
Video-based pedestrian reidentification is an emerging task in video surveillance and is closely related to several real-world applications. Its goal is to match pedestrians across multiple nonoverlapping network cameras. Despite the recent effort, the performance of pedestrian reidentification needs further improvement. Hence, we propose a novel two-stream multirate recurrent neural network for video-based pedestrian reidentification with two inherent advantages: First, capturing the static spatial and temporal information; Second,Author: Figure II is not cited in the text. Please cite it at the appropriate place. dealing with motion speed variance. Given video sequences of pedestrians, we start with extracting spatial and motion features using two different deep neural networks. Then, we explore the feature correlation which results in a regularized fusion network integrating the two aforementioned networks. Considering that pedestrians, sometimes even the same pedestrian, move in different speeds across different camera views, we extend our approach by feeding the two networks into a multirate recurrent network to exploit the temporal correlations. Extensive experiments have been conducted on two real-world video-based pedestrian reidentification benchmarks: iLIDS-VID and PRID 2011 datasets. The experimental results confirm the efficacy of the proposed method. Our code will be released upon acceptance.
Zhihui Li 0001, De Cheng, Huaxiang Zhang 0001, Kun Zhan, Yi Yang 0001
IEEE Trans. Ind. Informatics3
2017 Complex Event Detection by Identifying Reliable Shots from Untrimmed Videos
abstract
The goal of complex event detection is to automatically detect whether an event of interest happens in temporally untrimmed long videos which usually consist of multiple video shots. Observing some video shots in positive (resp. negative) videos are irrelevant (resp. relevant) to the given event class, we formulate this task as a multi-instance learning (MIL) problem by taking each video as a bag and the video shots in each video as instances. To this end, we propose a new MIL method, which simultaneously learns a linear SVM classifier and infers a binary indicator for each instance in order to select reliable training instances from each positive or negative bag. In our new objective function, we balance the weighted training errors and a l1-l2mixed-norm regularization term which adaptively selects reliable shots as training instances from different videos to have them as diverse as possible. We also develop an alternating optimization approach that can efficiently solve our proposed objective function. Extensive experiments on the challenging real-world Multimedia Event Detection (MED) datasets MEDTest-14, MEDTest-13 and CCV clearly demonstrate the effectiveness of our proposed MIL approach for complex event detection.
Hehe Fan, Xiaojun Chang, De Cheng, Yi Yang 0001, Dong Xu 0001, Alex Hauptmann 0001
ICCV3
2017 Rewind to track: Parallelized apprenticeship learning with backward tracklets
abstract
Data association, which could be categorized into offline approaches and the online counterparts, is a crucial part of a multi-object tracker in the tracking-by-detection framework. On the one hand, classical offline data association methods exploit all the video data and have high computation cost, which makes them unscalable to long-term offline video data. On the other hand, online approaches have much lower computation cost, but they suffer from ID-switches and tracklet drifting problem when directly applied to offline data as they are only aware of “past” observations. In this paper, we propose a mixed style tracker, which is not only as efficient as the online tracker but also aware of “future” observations in offline setting. We start from a Markov Decision Process (MDP) online tracker and design a parallelized apprenticeship learning algorithm to learn both the reward function and transition policy in MDP. By proposing a rewind to track strategy to generate backward tracklets, future detections in offline data are efficiently utilized to obtain a more stable similarity measurement for association. Experiment results show that our approach achieves the state-of-the-art performance on challenging datasets.
Jiang Liu 0011, Jia Chen 0001, De Cheng, Chenqiang Gao, Alex Hauptmann 0001
ICME3
2017 Discriminative Dictionary Learning With Ranking Metric Embedded for Person Re-Identification
abstract
The goal of person re-identification (Re-Id) is to match pedestrians captured from multiple non-overlapping cameras. In this paper, we propose a novel dictionary learning based method with the ranking metric embedded, for person Re-Id. A new and essential ranking graph Laplacian term is introduced, which minimizes the intra-personal compactness and maximizes the inter-personal dispersion in the objective. Different from the traditional dictionary learning based approaches and their extensions, which just use the same or not information, our proposed method can explore the ranking relationship among the person images, which is essential for such retrieval related tasks. Simultaneously, one distance measurement has been explicitly learned in the model to further improve the performance. Since we have reformulated these ranking constraints into the graph Laplacian form, the proposed method is easy-to-implement but effective. We conduct extensive experiments on three widely used person Re-Id benchmark datasets, and achieve state-of-the-art performances.
De Cheng, Xiaojun Chang, Li Liu 0031, Alex Hauptmann 0001, Yihong Gong, Nanning Zheng 0001
IJCAI1
2017 Video Search via Ranking Network with Very Few Query Exemplars
De Cheng, Lu Jiang 0004, Yihong Gong, Nanning Zheng 0001, Alex Hauptmann 0001
MMM (2)1
2017 Part-aware trajectories association across non-overlapping uncalibrated cameras
De Cheng, Yihong Gong, Jinjun Wang, Qiqi Hou, Nanning Zheng 0001
Neurocomputing1
2017 A Weight-Adaptive Laplacian Embedding for Graph-Based Clustering
abstract
Graph-based clustering methods perform clustering on a fixed input data graph. Thus such clustering results are sensitive to the particular graph construction. If this initial construction is of low quality, the resulting clustering may also be of low quality. We address this drawback by allowing the data graph itself to be adaptively adjusted in the clustering procedure. In particular, our proposed weight adaptive Laplacian (WAL) method learns a new data similarity matrix that can adaptively adjust the initial graph according to the similarity weight in the input data graph. We develop three versions of these methods based on the L2-norm, fuzzy entropy regularizer, and another exponential-based weight strategy, that yield three new graph-based clustering objectives. We derive optimization algorithms to solve these objectives. Experimental results on synthetic data sets and real-world benchmark data sets exhibit the effectiveness of these new graph-based clustering methods.
De Cheng, Feiping Nie 0001, Jiande Sun 0001, Yihong Gong
Neural Comput.1
2017 Balanced Mixture of Deformable Part Models With Automatic Part Configurations
abstract
This paper presents a method to improve the traditional mixture of deformable part models (MDPM) method from the learning perspective. First, an object part configuration learning algorithm based on group sparsity constraint is introduced to automatically discover the object part number, size, and location. The algorithm imposes two additional regularization terms in addition to the standard hinge loss function. The first term focuses on automatic part selection and the second term focuses on automatic part placement. Second, this paper introduces an improved MDPM training framework. The framework applies a learned transformation to normalize the prediction score from each individual deformable part model (DPM) into a pseudo probability such that the partition of the entire object appearance feature space becomes less sensitive to the prior distributions of different DPMs. Finally, the two proposed improvements are combined and formulated under the expectation-maximization framework. We evaluate our method mainly using the PASCAL VOC2007 and VOC2010 detection benchmarks and show that the proposed learning algorithms could increase the detection mean AP score by 2.4% and 0.9%, respectively, on these two data sets when using the proposed part selection method and the training algorithm. We also present further in-depth analysis of the proposed algorithm in the experiments.
De Cheng, Yihong Gong, Jingjun Wang, Nanning Zheng 0001
IEEE Trans. Circuits Syst. Video Technol.1
2016 Person Re-identification by Multi-Channel Parts-Based CNN with Improved Triplet Loss Function
abstract
Person re-identification across cameras remains a very challenging problem, especially when there are no overlapping fields of view between cameras. In this paper, we present a novel multi-channel parts-based convolutional neural network (CNN) model under the triplet framework for person re-identification. Specifically, the proposed CNN model consists of multiple channels to jointly learn both the global full-body and local body-parts features of the input persons. The CNN model is trained by an improved triplet loss function that serves to pull the instances of the same person closer, and at the same time push the instances belonging to different persons farther from each other in the learned feature space. Extensive comparative evaluations demonstrate that our proposed method significantly outperforms many state-of-the-art approaches, including both traditional and deep network-based ones, on the challenging i-LIDS, VIPeR, PRID2011 and CUHK01 datasets.
De Cheng, Yihong Gong, Sanping Zhou, Jinjun Wang, Nanning Zheng 0001
CVPR1
2015 Training mixture of weighted SVM for object detection using EM algorithm
De Cheng, Jinjun Wang, Xing Wei 0001, Yihong Gong
Neurocomputing1