Chuchu Han

dblp:226/2692 · DBLP profile ↗
← Back
19ranked-venue papers
8as first author
8since 2021 · last 2025
0000-0001-9403-353XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 15 · 6 first-author · 5 since 2021Artificial intelligence and machine learning · 12 · 5 first-author · 7 since 2021
YearPublicationVenuePosition
2025 Spatial cascaded clustering and weighted memory for unsupervised person re-identification
Jiahao Hong, Jialong Zuo, Chuchu Han, Ruochen Zheng, Ming Tian, Changxin Gao, Nong Sang
Image Vis. Comput.3
2025 Linear Feature Source Prediction and Recombination Network for Noisy Label Learning
abstract
Collecting training data for deep models from the Internet is a common data acquisition approach. However, there are challenges in using these data directly, as they often contain inaccurate annotations. This situation has increased the attention and importance of noisy label learning, the process of training a deep model with unreliable annotations. The typical strategy in noisy label learning is to identify potential mislabeled samples and assign pseudo-labels generated by the network to them, replacing the original labels. However, existing methods encounter the following problems: 1) they typically do not evaluate the pseudo-labels and directly use all of them, and 2) empirical parameter settings are often dataset-specific. These shortcomings limit the application of these methods in real-world scenarios. In this paper, we propose the Linear Feature Source Prediction and Recombination Network (LFSPR), trying to solve the problem above by proposing a new pretext task. The pretext task is designed to build the linear connection between the high-dimensional feature and the low-dimensional feature. The source of the latter is regarded as the high-dimensional feature, which follows a non-linear head network to obtain the low-dimensional feature. The pretext task is designed in low-dimensional space by predicting the linear composition weights of the potential source. Based on the pretext task, our method can generate pseudo-labels for uncertain samples while dynamically evaluating and selecting them, rather than simply using all pseudo-labels or discarding a fixed proportion of pseudo-labels for a given dataset. To the best of our knowledge, this is the first approach in the noisy label learning domain to employ pretext task for the pseudo-labels generation, evaluation and selection. The experiments on CIFAR-10, CIFAR-100 and Clothing1M demonstrate the effectiveness of our method.
Ruochen Zheng, Chuchu Han, Changxin Gao, Nong Sang
IEEE Trans. Circuits Syst. Video Technol.2
2023 DMRNet++: Learning Discriminative Features With Decoupled Networks and Enriched Pairs for One-Step Person Search
abstract
Person search aims at localizing and recognizing query persons from raw video frames, which is a combination of two sub-tasks, i.e., pedestrian detection and person re-identification. The dominant fashion is termed as the one-step person search that jointly optimizes detection and identification in a unified network, exhibiting higher efficiency. However, there remain major challenges: (i) conflicting objectives of multiple sub-tasks under the shared feature space, (ii) inconsistent memory bank caused by the limited batch size, (iii) underutilized unlabeled identities during the identification learning. To address these issues, we develop an enhanced decoupled and memory-reinforced network (DMRNet++). First, we simplify the standard tightly coupled pipelines and establish a task-decoupled framework (TDF). Second, we build a memory-reinforced mechanism (MRM), with a slow-moving average of the network to better encode the consistency of the memorized features. Third, considering the potential of unlabeled samples, we model the recognition process as semi-supervised learning. An unlabeled-aided contrastive loss (UCL) is developed to boost the identification feature learning by exploiting the aggregation of unlabeled identities. Experimentally, the proposed DMRNet++ obtains the mAP of 94.5% and 52.1% on CUHK-SYSU and PRW datasets, which exceeds most existing methods.
Chuchu Han, Zhedong Zheng, Dongdong Yu, Zehuan Yuan, Changxin Gao, Nong Sang, Yi Yang 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2022 Multi-Centroid Representation Network for Domain Adaptive Person Re-ID
abstract
Recently, many approaches tackle the Unsupervised Domain Adaptive person re-identification (UDA re-ID) problem through pseudo-label-based contrastive learning. During training, a uni-centroid representation is obtained by simply averaging all the instance features from a cluster with the same pseudo label. However, a cluster may contain images with different identities (label noises) due to the imperfect clustering results, which makes the uni-centroid representation inappropriate. In this paper, we present a novel Multi-Centroid Memory (MCM) to adaptively capture different identity information within the cluster. MCM can effectively alleviate the issue of label noises by selecting proper positive/negative centroids for the query image. Moreover, we further propose two strategies to improve the contrastive learning process. First, we present a Domain-Specific Contrastive Learning (DSCL) mechanism to fully explore intra-domain information by comparing samples only from the same domain. Second, we propose Second-Order Nearest Interpolation (SONI) to obtain abundant and informative negative samples. We integrate MCM, DSCL, and SONI into a unified framework named Multi-Centroid Representation Network (MCRN). Extensive experiments demonstrate the superiority of MCRN over state-of-the-art approaches on multiple UDA re-ID tasks and fully unsupervised re-ID tasks.
Tengteng Huang, Chi Zhang 0026, Yuanjie Shao, Chuchu Han, Changxin Gao, Nong Sang
AAAI6
2022 Content-Variant Reference Image Quality Assessment via Knowledge Distillation
abstract
Generally, humans are more skilled at perceiving differences between high-quality (HQ) and low-quality (LQ) images than directly judging the quality of a single LQ image. This situation also applies to image quality assessment (IQA). Although recent no-reference (NR-IQA) methods have made great progress to predict image quality free from the reference image, they still have the potential to achieve better performance since HQ image information is not fully exploited. In contrast, full-reference (FR-IQA) methods tend to provide more reliable quality evaluation, but its practicability is affected by the requirement for pixel-level aligned reference images. To address this, we firstly propose the content-variant reference method via knowledge distillation (CVRKD-IQA). Specifically, we use non-aligned reference (NAR) images to introduce various prior distributions of high-quality images. The comparisons of distribution differences between HQ and LQ images can help our model better assess the image quality. Further, the knowledge distillation transfers more HQ-LQ distribution difference information from the FR-teacher to the NAR-student and stabilizing CVRKD-IQA performance. Moreover, to fully mine the local-global combined information, while achieving faster inference speed, our model directly processes multiple image patches from the input with the MLP-mixer. Cross-dataset experiments verify that our model can outperform all NAR/NR-IQA SOTAs, even reach comparable performance than FR-IQA methods on some occasions. Since the content-variant and non-aligned reference HQ images are easy to obtain, our model can support more IQA applications with its robustness to content variations. Our code is available: https://github.com/guanghaoyin/CVRKD-IQA.
Guanghao Yin, Wei Wang 0009, Zehuan Yuan, Chuchu Han, Wei Ji 0008, Shouqian Sun, Changhu Wang
AAAI4
2022 Single image based 3D human pose estimation via uncertainty learning
Chuchu Han, Xin Yu 0002, Changxin Gao, Nong Sang, Yi Yang 0001
Pattern Recognit.1
2021 Decoupled and Memory-Reinforced Networks: Towards Effective Feature Learning for One-Step Person Search
abstract
The goal of person search is to localize and match query persons from scene images. For high efficiency, one-step methods have been developed to jointly handle the pedestrian detection and identification sub-tasks using a single network. There are two major challenges in the current one-step approaches. One is the mutual interference between the optimization objectives of multiple sub-tasks. The other is the sub-optimal identification feature learning caused by small batch size when end-to-end training. To overcome these problems, we propose a decoupled and memory-reinforced network (DMRNet). Specifically, to reconcile the conflicts of multiple objectives, we simplify the standard tightly coupled pipelines and establish a deeply decoupled multi-task learning framework. Further, we build a memory-reinforced mechanism to boost the identification feature learning. By queuing the identification features of recently accessed instances into a memory bank, the mechanism augments the similarity pair construction for pairwise metric learning. For better encoding consistency of the stored features, a slow-moving average of the network is applied for extracting these features. In this way, the dual networks reinforce each other and converge to robust solution states. Experimentally, the proposed method obtains 93.2% and 46.9% mAP on CUHK-SYSU and PRW datasets, which exceeds all the existing one-step methods.
Chuchu Han, Zhedong Zheng, Changxin Gao, Nong Sang, Yi Yang 0001
AAAI1
2021 Weakly Supervised Person Search with Region Siamese Networks
abstract
Supervised learning is dominant in person search, but it requires elaborate labeling of bounding boxes and identities. Large-scale labeled training data is often difficult to collect, especially for person identities. A natural question is whether a good person search model can be trained without the need of identity supervision. In this paper, we present a weakly supervised setting where only bounding box annotations are available. Based on this new setting, we provide an effective baseline model termed Region Siamese Networks (R-SiamNets). Towards learning useful representations for recognition in the absence of identity labels, we supervise the R-SiamNet with instance-level consistency loss and cluster-level contrastive loss. For instance-level consistency learning, the R-SiamNet is constrained to extract consistent features from each person region with or without out-of-region context. For cluster-level contrastive learning, we enforce the aggregation of closest instances and the separation of dissimilar ones in feature space. Extensive experiments validate the utility of our weakly supervised method. Our model achieves the rank-1 of 87.1% and mAP of 86.0% on CUHK-SYSU benchmark, which surpasses several fully supervised methods, such as OIM [36] and MGTS [4], by a clear margin. More promising performance can be reached by incorporating extra training data. We hope this work could encourage the future research in this field.
Chuchu Han, Dongdong Yu, Zehuan Yuan, Changxin Gao, Nong Sang, Yi Yang 0001, Changhu Wang
ICCV1
2020 Deep Representation Learning on Long-Tailed Data: A Learnable Embedding Augmentation Perspective
abstract
This paper considers learning deep features from long-tailed data. We observe that in the deep feature space, the head classes and the tail classes present different distribution patterns. The head classes have a relatively large spatial span, while the tail classes have a significantly small spatial span, due to the lack of intra-class diversity. This uneven distribution between head and tail classes distorts the overall feature space, which compromises the discriminative ability of the learned features. In response, we seek to expand the distribution of the tail classes during training, so as to alleviate the distortion of the feature space. To this end, we propose to augment each instance of the tail classes with certain disturbances in the deep feature space. With the augmentation, a specified feature vector becomes a set of probable features scattered around itself, which is analogical to an atomic nucleus surrounded by the electron cloud. Intuitively, we name it as ``feature cloud''. The intra-class distribution of the feature cloud is learned from the head classes, and thus provides higher intra-class variation to the tail classes. Consequentially, it alleviates the distortion of the learned feature space, and improves deep representation learning on long tailed data. Extensive experimental evaluations on person re-identification and face recognition tasks confirm the effectiveness of our method.
Jialun Liu, Yifan Sun 0003, Chuchu Han, Zhaopeng Dou, Wenhui Li 0002
CVPR3
2020 Do Not Disturb Me: Person Re-identification Under the Interference of Other Pedestrians
Shizhen Zhao, Changxin Gao, Jun Zhang 0018, Hao Cheng 0012, Chuchu Han, Xinyang Jiang, Wei-Shi Zheng 0001, Nong Sang, Xing Sun 0001
ECCV (6)5
2020 Keypoint-Based Feature Matching For Partial Person Re-Identification
abstract
As a derivative of person re-identification (re-ID), partial re- ID aims to retrieve a partial pedestrian across holistic person images captured by non-overlapping cameras. This task is more challenging and closer to real-world applications. Since we cannot locate the part of the partial image, the misaligned region compromises the performance greatly when directly (a) compare a partial pedestrian with a holistic one. To alleviate this issue, we propose a Keypoint-Based Feature Matching (KBFM) network, which constructs a simple and effective framework for partial re-ID. Specifically, our architecture explicitly leverages the keypoints generated by pose estimation. Based on the visible keypoints, coordinates of the corresponding visible region can be computed. And the keypoint-based feature embeddings can be generated by bilinear sampling. When matching two images, we extract their features on the basis of shared visible keypoints, avoiding the misalignment and disturbance. Moreover, considering the triplet loss cannot be flexibly built in the partial re-ID pipeline, we improve the original sampling method and achieve significant performance. Extensive experimental results on two widely used benchmarks demonstrate significant performance improvements of our method over most state-of-the-art methods.
Chuchu Han, Changxin Gao, Nong Sang
ICIP1
2020 Hard sample mining makes person re-identification more efficient and accurate
Kezhou Chen, Chuchu Han, Nong Sang, Changxin Gao
Neurocomputing3
2020 Complementation-Reinforced Attention Network for Person Re-Identification
abstract
Fine-grained information has been proved helpful for person re-identification, and multi-head attention mechanism offers a feasible solution for this. However, we observe severe redundancy among the multiple branches, which might make the learned representation over-emphasize certain discriminative regions and correspondingly ignore other potentially informative regions. Therefore, we tackle this issue by two aspects yielding the so-called Complementation-Reinforced Attention Network (CRAN). One is the redundancy among branches, and we propose to impose complementing constraints among multiple attention heads. The constraints are two-fold: on the one hand, it encourages each branch to attend to complementary attention regions; on the other hand, it enforces orthogonality among the learned features of different regions in the embedding space. The other is the redundancy among query positions for each attention head. So we simplify the attention block by sparsifying the query positions. Besides, in order to achieve efficient retrieval, we propose an adaptive feature fusion method for dimensional reduction. Compared with the commonly used feature ensemble, our method effectively reduces the dimensionality while keeping the discriminative ability. We demonstrate the effectiveness of our method on MSMT17, Market-1501, DukeMTMC-reID, and CUHK03 datasets.
Chuchu Han, Ruochen Zheng, Changxin Gao, Nong Sang
IEEE Trans. Circuits Syst. Video Technol.1
2019 Camera Style and Identity Disentangling Network for Person Re-identification
Ruochen Zheng, Lerenhan Li, Chuchu Han, Changxin Gao, Nong Sang
BMVC3
2019 Re-ID Driven Localization Refinement for Person Search
abstract
Person search aims at localizing and identifying a query person from a gallery of uncropped scene images. Different from person re-identification (re-ID), its performance also depends on the localization accuracy of a pedestrian detector. The state-of-the-art methods train the detector individually, and the detected bounding boxes may be sub-optimal for the following re-ID task. To alleviate this issue, we propose a re-ID driven localization refinement framework for providing the refined detection boxes for person search. Specifically, we develop a differentiable ROI transform layer to effectively transform the bounding boxes from the original images. Thus, the box coordinates can be supervised by the re-ID training other than the original detection task. With this supervision, the detector can generate more reliable bounding boxes, and the downstream re-ID model can produce more discriminative embeddings based on the refined person localizations. Extensive experimental results on the widely used benchmarks demonstrate that our proposed method performs favorably against the state-of-the-art person search methods.
Chuchu Han, Jiacheng Ye, Yunshan Zhong, Xin Tan 0002, Chi Zhang 0026, Changxin Gao, Nong Sang
ICCV1
2018 Improving Person Re-Identification by Adaptive Hard Sample Mining
abstract
The field of person reidentification has made significant advances riding on the wave of deep learning. However, owing to the fact that there are much more easy examples than those meaningful hard examples in dataset, the training tends to stagnate quickly and the model may suffer from over-fitting. Therefore, the hard sample mining method is fateful to optimize the model and improve the learning efficiency. In this paper, an Adaptive Hard Sample Mining algorithm is proposed for training a robust person re-identification model. No need for hand-picking the images in the batch or designing the loss function for both positive and negative pairs, we can briefly calculate the hard level by comparing the prediction result with the true label of the sample. Meanwhile, an adaptive threshold of hard level can make the algorithm not only stay in step with training process harmoniously but also alleviate the under-fitting and over-fitting problem simultaneously. Besides, the designed network to implement the approach has good generalization performance that can be combined with various of existing models readily. Experimental results on Market-1501 and DukeMTMC-reID datasets clearly demonstrate the effectiveness of the proposed algorithm.
Kezhou Chen, Chuchu Han, Nong Sang, Changxin Gao, Ruolin Wang
ICIP3
2018 Re-ranking Person Re-identification with Adaptive Hard Sample Mining
Chuchu Han, Kezhou Chen, Jin Wang 0019, Changxin Gao, Nong Sang
PRCV (1)1
2018 Global Feature Learning with Human Body Region Guided for Person Re-identification
Nong Sang, Kezhou Chen, Chuchu Han, Changxin Gao
PRCV (1)4
2018 Center-Level Verification Model for Person Re-identification
Ruochen Zheng, Changqian Yu, Chuchu Han, Changxin Gao, Nong Sang
PRCV (1)4