Ancong Wu

dblp:168/9430 · DBLP profile ↗
← Back
48ranked-venue papers
9as first author
25since 2021 · last 2025
0000-0002-7969-3190ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 37 · 6 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 36 · 7 first-author · 17 since 2021
YearPublicationVenuePosition
2025 MaintaAvatar: A Maintainable Avatar Based on Neural Radiance Fields by Continual Learning
abstract
The generation of a virtual digital avatar is a crucial research topic in the field of computer vision. Many existing works utilize Neural Radiance Fields (NeRF) to address this issue and have achieved impressive results. However, previous works assume the images of the training person are available and fixed while the appearances and poses of a subject could constantly change and increase in real-world scenarios. How to update the human avatar but also maintain the ability to render the old appearance of the person is a practical challenge. One trivial solution is to combine the existing virtual avatar models based on NeRF with continual learning methods. However, there are some critical issues in this approach: learning new appearances and poses can cause the model to forget past information, which in turn leads to a degradation in the rendering quality of past appearances, especially color bleeding issues, and incorrect human body poses. In this work, we propose a maintainable avatar (MaintaAvatar) based on neural radiance fields by continual learning, which resolves the issues by utilizing a Global-Local Joint Storage Module and a Pose Distillation Module. Overall, our model requires only limited data collection to quickly fine-tune the model while avoiding catastrophic forgetting, thus achieving a maintainable virtual avatar. The experimental results validate the effectiveness of our MaintaAvatar model.
Shengbo Gu, Yu-Kun Qiu, Yu-Ming Tang, Ancong Wu, Wei-Shi Zheng 0001
AAAI4
2025 iManip: Skill-Incremental Learning for Robotic Manipulation
Zexin Zheng, Jia-Feng Cai, Xiao-Ming Wu 0002, Yi-Lin Wei, Yu-Ming Tang, Ancong Wu, Wei-Shi Zheng 0001
ICCV6
2025 Progressive Human Motion Generation Based on Text and Few Motion Frames
abstract
Although existing text-to-motion (T2M) methods can produce realistic human motion from text description, it is still difficult to align the generated motion with the desired postures since using text alone is insufficient for precisely describing diverse postures. To achieve more controllable generation, an intuitive way is to allow the user to input a few motion frames describing precise desired postures. Thus, we explore a new Text-Frame-to-Motion (TF2M) generation task that aims to generate motions from text and very few given frames. Intuitively, the closer a frame is to a given frame, the lower the uncertainty of this frame is when conditioned on this given frame. Hence, we propose a novel Progressive Motion Generation (PMG) method to progressively generate a motion from the frames with low uncertainty to those with high uncertainty in multiple stages. During each stage, new frames are generated by a Text-Frame Guided Generator conditioned on frame-aware semantics of the text, given frames, and frames generated in previous stages. Additionally, to alleviate the train-test gap caused by multi-stage accumulation of incorrectly generated frames during testing, we propose a Pseudo-frame Replacement Strategy for training. Experimental results show that our PMG outperforms existing T2M generation methods by a large margin with even one given frame, validating the effectiveness of our PMG. Code is available here.
Ling-An Zeng, Gaojie Wu, Ancong Wu, Jianfang Hu, Wei-Shi Zheng 0001
IEEE Trans. Circuits Syst. Video Technol.3
2024 Factorized Diffusion Autoencoder for Unsupervised Disentangled Representation Learning
abstract
Unsupervised disentangled representation learning aims to recover semantically meaningful factors from real-world data without supervision, which is significant for model generalization and interpretability. Current methods mainly rely on assumptions of independence or informativeness of factors, regardless of interpretability. Intuitively, visually interpretable concepts better align with human-defined factors. However, exploiting visual interpretability as inductive bias is still under-explored. Inspired by the observation that most explanatory image factors can be represented by ``content + mask'', we propose a content-mask factorization network (CMFNet) to decompose an image into different groups of content codes and masks, which are further combined as content masks to represent different visual concepts. To ensure informativeness of the representations, the CMFNet is jointly learned with a generator conditioned on the content masks for reconstructing the input image. The conditional generator employs a diffusion model to leverage its robust distribution modeling capability. Our model is called the Factorized Diffusion Autoencoder (FDAE). To enhance disentanglement of visual concepts, we propose a content decorrelation loss and a mask entropy loss to decorrelate content masks in latent space and spatial space, respectively. Experiments on Shapes3d, MPI3D and Cars3d show that our method achieves advanced performance and can generate visually interpretable concept-specific masks. Source code and supplementary materials are available at https://github.com/wuancong/FDAE.
Ancong Wu, Wei-Shi Zheng 0001
AAAI1
2024 Bridge Past and Future: Overcoming Information Asymmetry in Incremental Object Detection
Qijie Mo, Yipeng Gao, Shenghao Fu, Junkai Yan, Ancong Wu, Wei-Shi Zheng 0001
ECCV (16)5
2024 DreamView: Injecting View-Specific Text Guidance Into Text-to-3D Generation
Junkai Yan, Yipeng Gao, Qize Yang, Xihan Wei, Xuansong Xie, Ancong Wu, Wei-Shi Zheng 0001
ECCV (25)6
2024 Privacy-Preserving Face Recognition with Adaptive Generative Perturbations
Delong Zhang, Yixing Peng, Ancong Wu, Wei-Shi Zheng 0001
ICPR (14)3
2024 Fine-Grained Depth Knowledge Distillation for Cloth-Changing Person Re-identification
abstract
The mission of cloth-changing person re-identification (CC-ReID) is to discover cloth-invariant and identity-related cues, while traditional person ReID methods rely on appearance features that are biased to cloth-related cues. To tackle this cloth-biased problem, many CC-ReID methods introduced auxiliary body shape information to extract cloth-invariant features, such as 2D sketch images or the 3D Skinned Multi-Person Linear (SMPL) model. However, 2D auxiliary information lacks 3D spatial features, while the 3D SMPL model encounters challenges in capturing features at a finer granularity due to manually defined parameters. To extract fine-grained 3D shape features, we estimate depth maps that contain richer shape information and propose a Fine-grained Depth feature Mining and Distillation (FDMD) framework. We introduce a depth branch and design a fine-grained local feature interaction module to mine fine-grained 3D body shape knowledge from estimated depth maps by exploring the context of semantic-aware local body-part features. To integrate cloth-invariant depth knowledge into the appearance features, the fine-grained 3D shape features are transferred to an appearance branch by feature-space-aligned distillation. Extensive experiments demonstrate that FDMD can achieve state-of-the-art performance on three widely used CC-ReID benchmarks PRCC, Celeb-reID and LaST.
Yuhan Yao 0002, Ancong Wu, Jiangqun Ni, Wei-Shi Zheng 0001
IJCNN3
2024 PixelFade: Privacy-preserving Person Re-identification with Noise-guided Progressive Replacement
Delong Zhang, Yi-Xing Peng, Xiao-Ming Wu 0002, Ancong Wu, Wei-Shi Zheng 0001
ACM Multimedia4
2024 Asymmetric Mutual Learning for Unsupervised Transferable Visible-Infrared Re-Identification
abstract
Visible-infrared person re-identification (Re-ID) plays a crucial role in matching people across camera views in the darkness and normal lighting. To reduce annotation cost, it is advantageous to learn Re-ID model from unlabeled visible-infrared image pairs. However, large modality gap makes it difficult to discover the underlying cross-modality sample relations. Compared with cross-modality sample pairs in the target domain, it is easier to obtain more single-modality visible image samples from other domains. In this work, we study unsupervised transfer learning to extract modality-shared knowledge from auxiliary unlabeled visible images in a source domain and leverage this knowledge to learn cross-modality matching in the unlabeled target domain. Our framework consists of two stages: RGB-gray asymmetric mutual learning and unsupervised cross-modality self-training. In the first stage, to extract visible-infrared shared information from auxiliary unlabeled visible images, we regard RGB images and grayscale fake infrared images transformed from RGB images as two views to learn view-shared information and simultaneously preserve RGB-specific information. Based on information theoretic analysis, we learn an RGB-gray feature extractor and further introduce an auxiliary gray feature extractor to quantify RGB-gray shared knowledge. This knowledge is then transferred to the RGB-gray feature extractor without eliminating RGB-specific information. We call this process Cross-Modality Asymmetric Mutual Learning (CMAM). In the second stage, for unsupervised cross-modality self-training in the target domain, we fuse the complementary knowledge in two models by mutual learning and employ bipartite cross-modality pseudo labeling to alleviate modality gap. For a more extensive evaluation, we collected a new public multi-modality dataset, SYSU-MM02, constructed from untrimmed videos. Our method achieves the state-of-the-art performance on three benchmark datasets.
Ancong Wu, Chengzhi Lin, Wei-Shi Zheng 0001
IEEE Trans. Circuits Syst. Video Technol.1
2024 Online Multi-View Learning With Knowledge Registration Units
abstract
In this work, we investigate online multi-view learning according to the multi-view complementarity and consistency principles to memorably process online multi-view data when fused across views. Online diverse features through different deep feature extractors under different views are used as input to an online learning method to privately and memorably optimize in each view for the discovery and memorization of the view-specific information. More specifically, according to the multi-view complementarity principle, a softmax-weighted reducible (SWR) loss is proposed to selectively retain credible views and neglect incredible ones for the online model's cross-view complementarity fusion. According to the multi-view consistency principle, we design a cross-view embedding consistency (CVEC) loss and a cross-view Kullback-Leibler (CVKL) divergence loss to maintain the cross-view consistency of the online model. Since the online multi-view learning setup needs to avoid repeatedly accessing online data to handle the knowledge forgetting in each view, we propose a knowledge registration unit (KRU) based on dictionary learning to incrementally register newly view-specific knowledge of online unlabeled data to the learnable and adjustable dictionary. Finally, by using the above strategies, we propose an online multi-view KRU approach and evaluate it with comprehensive experiments, thereby showing its superiority in online multi-view learning.
Ancong Wu, Wei-Shi Zheng 0001
IEEE Trans. Neural Networks Learn. Syst.2
2023 Shape-Erased Feature Learning for Visible-Infrared Person Re-Identification
abstract
Due to the modality gap between visible and infrared images with high visual ambiguity, learning diverse modality-shared semantic concepts for visible-infrared person re-identification (VI-ReID) remains a challenging problem. Body shape is one of the significant modality-shared cues for VI-ReID. To dig more diverse modality-shared cues, we expect that erasing body-shape-related semantic concepts in the learned features can force the ReID model to extract more and other modality-shared features for identification. To this end, we propose shape-erased feature learning paradigm that decorrelates modality-shared features in two orthogonal subspaces. Jointly learning shape-related feature in one subspace and shape-erased features in the orthogonal complement achieves a conditional mutual information maximization between shape-erased feature and identity discarding body shape information, thus enhancing the diversity of the learned representation explicitly. Extensive experiments on SYSU-MM01, RegDB, and HITSZ-VCM datasets demonstrate the effectiveness of our method.
Ancong Wu, Wei-Shi Zheng 0001
CVPR2
2023 Rewarded Semi-Supervised Re-Identification on Identities Rarely Crossing Camera Views
abstract
Semi-supervised person re-identification (Re-ID) is an important approach for alleviating annotation costs when learning to match person images across camera views. Most existing works assume that training data contains abundant identities crossing camera views. However, this assumption is not true in many real-world applications, especially when images are captured in nonadjacent scenes for Re-ID in wider areas, where the identities rarely cross camera views. In this work, we operate semi-supervised Re-ID under a relaxed assumption of identities rarely crossing camera views, which is still largely ignored in existing methods. Since the identities rarely cross camera views, the underlying sample relations across camera views become much more uncertain, and deteriorate the noise accumulation problem in many advanced Re-ID methods that apply pseudo labeling for associating visually similar samples. To quantify such uncertainty, we parameterize the probabilistic relations between samples in a relation discovery objective for pseudo label training. Then, we introduce reward quantified by identification performance on a few labeled data to guide learning dynamic relations between samples for reducing uncertainty. Our strategy is called the Rewarded Relation Discovery (R$^{2}$D), of which the rewarded learning paradigm is under-explored in existing pseudo labeling methods. To further reduce the uncertainty in sample relations, we perform multiple relation discovery objectives learning to discover probabilistic relations based on different prior knowledge of intra-camera affinity and cross-camera style variation, and fuse the complementary knowledge of different probabilistic relations by similarity distillation. To better evaluate semi-supervised Re-ID on identities rarely crossing camera views, we collect a new real-world dataset called REID-CBD, and perform simulation on benchmark datasets. Experiment results show that our method outperforms a wide range of semi-supervised and unsupervised learning methods.
Ancong Wu, Wenhang Ge, Wei-Shi Zheng 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2022 Lifelong Person Re-identification by Pseudo Task Knowledge Preservation
abstract
In real world, training data for person re-identification (Re-ID) is collected discretely with spatial and temporal variations, which requires a model to incrementally learn new knowledge without forgetting old knowledge. This problem is called lifelong person re-identification (LReID). Variations of illumination and background for images of each task exhibit task-specific image style and lead to task-wise domain gap. In addition to missing data from the old tasks, task-wise domain gap is a key factor for catastrophic forgetting in LReID, which is ignored in existing approaches for LReID. The model tends to learn task-specific knowledge with task-wise domain gap, which results in stability and plasticity dilemma. To overcome this problem, we cast LReID as a domain adaptation problem and propose a pseudo task knowledge preservation framework to alleviate the domain gap. Our framework is based on a pseudo task transformation module which maps the features of the new task into the feature space of the old tasks to complement the limited saved exemplars of the old tasks. With extra transformed features in the task-specific feature space, we propose a task-specific domain consistency loss to implicitly alleviate the task-wise domain gap for learning task-shared knowledge instead of task-specific one. Furthermore, to guide knowledge preservation with the feature distributions of the old tasks, we propose to preserve knowledge on extra pseudo tasks which jointly distills knowledge and discriminates identity, in order to achieve a better trade-off between stability and plasticity for lifelong learning with task-wise domain gap. Extensive experiments demonstrate the superiority of our method as compared with the state-of-the-art lifelong learning and LReID methods.
Wenhang Ge, Junlong Du, Ancong Wu, Yuqiao Xian, Feiyue Huang, Wei-Shi Zheng 0001
AAAI3
2022 Camera-Conditioned Stable Feature Generation for Isolated Camera Supervised Person Re-IDentification
abstract
To learn camera-view invariant features for person Re-IDentification (Re-ID), the cross-camera image pairs of each person play an important role. However, such cross-view training samples could be unavailable under the ISo-lated Camera Supervised (ISCS) setting, e.g., a surveillance system deployed across distant scenes. To handle this challenging problem, a new pipeline is introduced by synthesizing the cross-camera samples in the feature space for model training. Specifically, the feature encoder and generator are end-to-end optimized under a novel method, Camera-Conditioned Stable Feature Generation (CCSFG). Its joint learning procedure raises concern on the stability of generative model training. Therefore, a new feature generator, σ-Regularized Conditional Variational Autoencoder (σ-Reg. CVAE), is proposed with theoretical and experimental analysis on its robustness. Extensive experiments on two ISCS person Re-ID datasets demonstrate the superiority of our CCSFG to the competitors.11https://github.com/ftd-Wuchao/CCSFG
Wenhang Ge, Ancong Wu, Xiaobin Chang
CVPR3
2022 Learning Multi-Context Dynamic Listwise Relation for Generalizable Person Re-Identification
abstract
Although Person re-identification (Re-ID) has made rapid development in supervised learning and domain adaptation, it is more desirable to learn a generalizable model that can be directly applied to unseen scenes without updating. Generalizable Re-ID is challenging due to uncertain cross-camera variations in unseen target domain, such as illumination and viewpoint change, which result in visual ambiguities. Existing generalizable Re-ID methods focus on learning more generalizable features for individual instances. They ignore the context in the ranking list of the target domain. When human encounters visual ambiguities when matching pedestrians in unfamiliar scenes, comparing similar instances in the ranking list and comparing environments in different cameras can help remove ambiguities and refine matching results. This is actually exploiting contextual information in target domain, which is ignored by existing generalizable Re-ID methods. For learning contextual information to refine matching in unseen target domain, we propose a Multi-Context Dynamic Listwise Relation Network (MDLRN) to extract and aggregate instance-level and camera-level contextual features of a list of images, which can dynamically adapt the metric to unseen cross-domain scene variations. We further propose camera-specific feature perturbation (CFP) to simulate cross-camera variations in unseen target domain to improve generalization. Extensive experiments showed the superiority of our method in domain generalization.
Chengzhi Lin, Ancong Wu, Wei-Shi Zheng 0001
ICPR2
2022 Text-Adaptive Multiple Visual Prototype Matching for Video-Text Retrieval
abstract
Cross-modal retrieval between videos and texts has gained increasing interest because of the rapid emergence of videos on the web. Generally, a video contains rich instance and event information and the query text only describes a part of the information. Thus, a video can have multiple different text descriptions and queries. We call it the Video-Text Correspondence Ambiguity problem. Current techniques mostly concentrate on mining local or multi-level alignment between contents of video and text (e.g., object to entity and action to verb). It is difficult for these methods to alleviate video-text correspondence ambiguity by describing a video using only one feature, which is required to be matched with multiple different text features at the same time. To address this problem, we propose a Text-Adaptive Multiple Visual Prototype Matching Model. It automatically captures multiple prototypes to describe a video by adaptive aggregation on video token features. Given a query text, the similarity is determined by the most similar prototype to find correspondence in the video, which is called text-adaptive matching. To learn diverse prototypes for representing the rich information in videos, we propose a variance loss to encourage different prototypes to attend to different contents of the video. Our method outperforms the state-of-the-art methods on four public video retrieval datasets.
Chengzhi Lin, Ancong Wu, Junwei Liang 0001, Jun Zhang 0018, Wenhang Ge, Wei-Shi Zheng 0001, Chunhua Shen
NeurIPS2
2022 Joint Bilateral-Resolution Identity Modeling for Cross-Resolution Person Re-Identification
Wei-Shi Zheng 0001, Jincheng Hong, Jiening Jiao, Ancong Wu, Xiatian Zhu, Shaogang Gong, Jiayin Qin, Jian-Huang Lai
Int. J. Comput. Vis.4
2021 One for More: Selecting Generalizable Samples for Generalizable ReID Model
abstract
Current training objectives of existing person Re-IDentification (ReID) models only ensure that the loss of the model decreases on selected training batch, with no regards to the performance on samples outside the batch. It will inevitably cause the model to over-fit the data in the dominant position (e.g., head data in imbalanced class, easy samples or noisy samples). The latest resampling methods address the issue by designing specific criterion to select specific samples that trains the model generalize more on certain type of data (e.g., hard samples, tail data), which is not adaptive to the inconsistent real world ReID data distributions. Therefore, instead of simply presuming on what samples are generalizable, this paper proposes a one-for-more training objective that directly takes the generalization ability of selected samples as a loss function and learn a sampler to automatically select generalizable samples. More importantly, our proposed one-for-more based sampler can be seamlessly integrated into the ReID training framework which is able to simultaneously train ReID models and the sampler in an end-to-end fashion. The experimental results show that our method can effectively improve the ReID model training and boost the performance of ReID models.
Enwei Zhang, Xinyang Jiang, Hao Cheng 0012, Ancong Wu, Fufu Yu, Ke Li 0015, Feng Zheng 0001, Wei-Shi Zheng 0001, Xing Sun 0001
AAAI4
2021 Fine-Grained Shape-Appearance Mutual Learning for Cloth-Changing Person Re-Identification
abstract
Recently, person re-identification (Re-ID) has achieved great progress. However, current methods largely depend on color appearance, which is not reliable when a person changes the clothes. Cloth-changing Re-ID is challenging since pedestrian images with clothes change exhibit large intra-class variation and small inter-class variation. Some significant features for identification are embedded in unobvious body shape differences across pedestrians. To explore such body shape cues for cloth-changing Re-ID, we propose a Fine-grained Shape-Appearance Mutual learning framework (FSAM), a two-stream framework that learns fine-grained discriminative body shape knowledge in a shape stream and transfers it to an appearance stream to complement the cloth-unrelated knowledge in the appearance features. Specifically, in the shape stream, FSAM learns fine-grained discriminative mask with the guidance of identities and extracts fine-grained body shape features by a pose-specific multi-branch network. To complement cloth-unrelated shape knowledge in the appearance stream, dense interactive mutual learning is performed across low-level and high-level features to transfer knowledge from shape stream to appearance stream, which enables the appearance stream to be deployed independently without extra computation for mask estimation. We evaluated our method on benchmark cloth-changing Re-ID datasets and achieved the start-of-the-art performance.
Peixian Hong, Ancong Wu, Xintong Han, Wei-Shi Zheng 0001
CVPR3
2021 Cross-Scene Person Trajectory Anomaly Detection Based on Re-Identification
abstract
In this work, we consider the cross-scene person trajectory anomaly detection problem, which detects the anomalous trajectories across multiple nonoverlapping scenes. This problem is highly significant for public security, but it is still underexplored. Since the trajectory is not continuous across nonoverlapping camera views, we take use of person reidentification (re-ID) to associate the same pedestrian in different scenes while mitigating its inaccuracy by a directional probabilistic graph. To better distinguishing normal samples from anomalies, We formulate a maximized margin graph autoencoder (MMGAE) model, and the reconstruction error of the MMGAE is regarded as an anomaly indicator for the sample. To verify the effectiveness of our approach, we collected and labeled a new dataset. we also explore the impact of the re-ID performance on the anomaly detection problem and the effect of an inaccurately constructed graph on the MMGAE.
Yuanxun Li, Ancong Wu, Wei-Shi Zheng 0001
ICME2
2021 Cross-Camera Feature Prediction for Intra-Camera Supervised Person Re-identification across Distant Scenes
abstract
Person re-identification (Re-ID) aims to match person images across non-overlapping camera views. The majority of Re-ID methods focus on small-scale surveillance systems in which each pedestrian is captured in different camera views of adjacent scenes. However, in large-scale surveillance systems that cover larger areas, it is required to track a pedestrian of interest across distant scenes (e.g., a criminal suspect escapes from one city to another). Since most pedestrians appear in limited local areas, it is difficult to collect training data with cross-camera pairs of the same person. In this work, we study intra-camera supervised person re-identification across distant scenes (ICS-DS Re-ID), which uses cross-camera unpaired data with intra-camera identity labels for training. It is challenging as cross-camera paired data plays a crucial role for learning camera-invariant features in most existing Re-ID methods. To learn camera-invariant representation from cross-camera unpaired training data, we propose a cross-camera feature prediction method to mine cross-camera self supervision information from camera-specific feature distribution by transforming fake cross-camera positive feature pairs and minimize the distances of the fake pairs. Furthermore, we automatically localize and extract local-level feature by a transformer. Joint learning of global-level and local-level features forms a global-local cross-camera feature prediction scheme for mining fine-grained cross-camera self supervision information. Finally, cross-camera self supervision and intra-camera supervision are aggregated in a framework. The experiments are conducted in the ICS-DS setting on Market-SCT, Duke-SCT and MSMT17-SCT datasets. The evaluation results demonstrate the superiority of our method, which gains significant improvements of 15.4 Rank-1 and 22.3 mAP on Market-SCT as compared to the second best method. Our code is available at https://github.com/g3956/CCFP.
Wenhang Ge, Chunyan Pan, Ancong Wu, Hongwei Zheng 0002, Wei-Shi Zheng 0001
ACM Multimedia3
2021 Letter-Level Online Writer Identification
Zelin Chen, Hong-Xing Yu, Ancong Wu, Wei-Shi Zheng 0001
Int. J. Comput. Vis.3
2021 Person Re-Identification by Contour Sketch Under Moderate Clothing Change
abstract
Person re-identification (re-id), the process of matching pedestrian images across different camera views, is an important task in visual surveillance. Substantial development of re-id has recently been observed, and the majority of existing models are largely dependent on color appearance and assume that pedestrians do not change their clothes across camera views. This limitation, however, can be an issue for re-id when tracking a person at different places and at different time if that person (e.g., a criminal suspect) changes his/her clothes, causing most existing methods to fail, since they are heavily relying on color appearance, and thus, they are inclined to match a person to another person wearing similar clothes. In this work, we call the person re-id under clothing change the "cross-clothes person re-id." In particular, we consider the case when a person only changes his clothes moderately as a first attempt at solving this problem based on visible light images; that is, we assume that a person wears clothes of a similar thickness, and thus the shape of a person would not change significantly when the weather does not change substantially within a short period of time. We perform cross-clothes person re-id based on a contour sketch of person image to take advantage of the shape of the human body instead of color information for extracting features that are robust to moderate clothing change. To select/sample more reliable and discriminative curve patterns on a body contour sketch, we introduce a learning-based spatial polar transformation (SPT) layer in the deep neural network to transform contour sketch images for extracting reliable and discriminant convolutional neural network (CNN) features in a polar coordinate space. An angle-specific extractor (ASE) is applied in the following layers to extract more fine-grained discriminant angle-specific features. By varying the sampling range of the SPT, we develop a multistream network for aggregating multi-granularity features to better identify a person. Due to the lack of a large-scale dataset for cross-clothes person re-id, we contribute a new dataset that consists of 33,698 images from 221 identities. Our experiments illustrate the challenges of cross-clothes person re-id and demonstrate the effectiveness of our proposed method.
Qize Yang, Ancong Wu, Wei-Shi Zheng 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2021 Online deep transferable dictionary learning
Ancong Wu, Wei-Shi Zheng 0001
Pattern Recognit.2
2020 Semi-supervised Person Re-identification by Attribute Similarity Guidance
abstract
Although supervised person re-identification (RE-ID) has achieved great progress with deep learning, it requires time-consuming annotation of a large number of pedestrian identities. To reduce labeling cost, we attempt to reduce cross-camera identity annotations and exploit pedestrian attribute annotations as auxiliary information instead. The pedestrian attributes, such as outfit styles, contain coarse semantic knowledge. Although pedestrian attributes are annotated without exhaustive searching in a camera network, which is much easier than cross-camera identity annotation, ambiguity exists in attributes when different persons have similar outfits. To solve this problem, we propose an Attribute Similarity Guidance loss (ASG) to guide appearance feature learning for RE-ID by selective attribute similarity preservation to avoid the impact of such ambiguity. Finally, we develop an attribute-guided self training framework to jointly utilize attribute annotations, unlabeled data and limited labeled data for semi-supervised learning. Extensive experiments on Market-1501 and DukeMTMC-ReID show the superiority of our method for semi-supervised RE-ID.
Peixian Hong, Ancong Wu, Wei-Shi Zheng 0001
ICPR2
2020 Transductive Multi-Object Tracking in Complex Events by Interactive Self-Training
abstract
Recently, multi-object tracking (MOT) for estimating trajectories of pedestrians has undergone fast development and played an important role in human-centric video analysis. However, video analysis in complex events (e.g. scenes in HiEve dataset) is still under-explored. In complex real-world scenarios, domain gap in unseen testing scenes and severe occlusion problem that disconnects tracks are challenging for existing online MOT methods without domain adaptation. To alleviate domain gap, we study the problem in a transductive learning setting, which assumes that unlabeled testing data is available for learning offline tracking. We propose a transductive interactive self-training method to adapt the tracking model to unseen crowded scenes with unlabeled testing data by means of teacher-student interative learning. To reduce prediction variance in an unseen domain, we train two different models and teach one model with pseudo labels of unlabeled data predicted by the other model interactively. To improve robustness against occlusions during self-training, we exploit disconnected track interpolation (DTI) to refine the predicted pseudo labels. Our method achieved MOTA of 60.23 on HiEve dataset and won the first place of Multi-person Motion Tracking in Complex Events (with Private Detection) in the ACM MM Grand Challenge on Large-scale Human-centric Video Analysis in Complex Events.
Ancong Wu, Chengzhi Lin, Bogao Chen, Weihao Huang, Wei-Shi Zheng 0001
ACM Multimedia1
2020 RGB-IR Person Re-identification by Cross-Modality Similarity Preservation
Ancong Wu, Wei-Shi Zheng 0001, Shaogang Gong, Jian-Huang Lai
Int. J. Comput. Vis.1
2020 Fine-Grained Person Re-identification
Jiahang Yin, Ancong Wu, Wei-Shi Zheng 0001
Int. J. Comput. Vis.2
2020 Unsupervised Person Re-Identification by Deep Asymmetric Metric Embedding
abstract
Person re-identification (Re-ID) aims to match identities across non-overlapping camera views. Researchers have proposed many supervised Re-ID models which require quantities of cross-view pairwise labelled data. This limits their scalabilities to many applications where a large amount of data from multiple disjoint camera views is available but unlabelled. Although some unsupervised Re-ID models have been proposed to address the scalability problem, they often suffer from the view-specific bias problem which is caused by dramatic variances across different camera views, e.g., different illumination, viewpoints and occlusion. The dramatic variances induce specific feature distortions in different camera views, which can be very disturbing in finding cross-view discriminative information for Re-ID in the unsupervised scenarios, since no label information is available to help alleviate the bias. We propose to explicitly address this problem by learning an unsupervised asymmetric distance metric based on cross-view clustering. The asymmetric distance metric allows specific feature transformations for each camera view to tackle the specific feature distortions. We then design a novel unsupervised loss function to embed the asymmetric metric into a deep neural network, and therefore develop a novel unsupervised deep framework named the DEep Clustering-based Asymmetric MEtric Learning (DECAMEL). In such a way, DECAMEL jointly learns the feature representation and the unsupervised asymmetric metric. DECAMEL learns a compact cross-view cluster structure of Re-ID data, and thus help alleviate the view-specific bias and facilitate mining the potential cross-view discriminative information for unsupervised Re-ID. Extensive experiments on seven benchmark datasets whose sizes span several orders show the effectiveness of our framework.
Hong-Xing Yu, Ancong Wu, Wei-Shi Zheng 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2019 Distilled Person Re-Identification: Towards a More Scalable System
abstract
Person re-identification (Re-ID), for matching pedestrians across non-overlapping camera views, has made great progress in supervised learning with abundant labelled data. However, the scalability problem is the bottleneck for applications in large-scale systems. We consider the scalability problem of Re-ID from three aspects: (1) low labelling cost by reducing label amount, (2) low extension cost by reusing existing knowledge and (3) low testing computation cost by using lightweight models. The requirements render scalable Re-ID a challenging problem. To solve these problems in a unified system, we propose a Multi-teacher Adaptive Similarity Distillation Framework, which requires only a few labelled identities of target domain to transfer knowledge from multiple teacher models to a user-specified lightweight student model without accessing source domain data. We propose the Log-Euclidean Similarity Distillation Loss for Re-ID and further integrate the Adaptive Knowledge Aggregator to select effective teacher models to transfer target-adaptive knowledge. Extensive evaluations show that our method can extend with high scalability and the performance is comparable to the state-of-the-art unsupervised and semi-supervised Re-ID methods.
Ancong Wu, Wei-Shi Zheng 0001, Jian-Huang Lai
CVPR1
2019 Patch-Based Discriminative Feature Learning for Unsupervised Person Re-Identification
abstract
While discriminative local features have been shown effective in solving the person re-identification problem, they are limited to be trained on fully pairwise labelled data which is expensive to obtain. In this work, we overcome this problem by proposing a patch-based unsupervised learning framework in order to learn discriminative feature from patches instead of the whole images. The patch-based learning leverages similarity between patches to learn a discriminative model. Specifically, we develop a PatchNet to select patches from the feature map and learn discriminative features for these patches. To provide effective guidance for the PatchNet to learn discriminative patch feature on unlabeled datasets, we propose an unsupervised patch-based discriminative feature learning loss. In addition, we design an image-level feature learning loss to leverage all the patch features of the same image to serve as an image-level guidance for the PatchNet. Extensive experiments validate the superiority of our method for unsupervised person re-id. Our code is available at https://github.com/QizeYang/PAUL.
Qize Yang, Hong-Xing Yu, Ancong Wu, Wei-Shi Zheng 0001
CVPR3
2019 Unsupervised Person Re-Identification by Soft Multilabel Learning
abstract
Although unsupervised person re-identification (RE-ID) has drawn increasing research attentions due to its potential to address the scalability problem of supervised RE-ID models, it is very challenging to learn discriminative information in the absence of pairwise labels across disjoint camera views. To overcome this problem, we propose a deep model for the soft multilabel learning for unsupervised RE-ID. The idea is to learn a soft multilabel (real-valued label likelihood vector) for each unlabeled person by comparing the unlabeled person with a set of known reference persons from an auxiliary domain. We propose the soft multilabel-guided hard negative mining to learn a discriminative embedding for the unlabeled target domain by exploring the similarity consistency of the visual features and the soft multilabels of unlabeled target pairs. Since most target pairs are cross-view pairs, we develop the cross-view consistent soft multilabel learning to achieve the learning goal that the soft multilabels are consistently good across different camera views. To enable effecient soft multilabel learning, we introduce the reference agent learning to represent each reference person by a reference agent in a joint embedding. We evaluate our unified deep model on Market-1501 and DukeMTMC-reID. Our model outperforms the state-of-the-art unsupervised RE-ID methods by clear margins. Code is available at https://github.com/KovenYu/MAR.
Hong-Xing Yu, Wei-Shi Zheng 0001, Ancong Wu, Shaogang Gong, Jian-Huang Lai
CVPR3
2019 Unsupervised Person Re-Identification by Camera-Aware Similarity Consistency Learning
abstract
For matching pedestrians across disjoint camera views in surveillance, person re-identification (Re-ID) has made great progress in supervised learning. However, it is infeasible to label data in a number of new scenes when extending a Re-ID system. Thus, studying unsupervised learning for Re-ID is important for saving labelling cost. Yet, cross-camera scene variation is a key challenge for unsupervised Re-ID, such as illumination, background and viewpoint variations, which cause domain shift in the feature space and result in inconsistent pairwise similarity distributions that degrade matching performance. To alleviate the effect of cross-camera scene variation, we propose a Camera-Aware Similarity Consistency Loss to learn consistent pairwise similarity distributions for intra-camera matching and cross-camera matching. To avoid learning ineffective knowledge in consistency learning, we preserve the prior common knowledge of intra-camera matching in the pretrained model as reliable guiding information, which does not suffer from cross-camera scene variation as cross-camera matching. To learn similarity consistency more effectively, we further develop a coarse-to-fine consistency learning scheme to learn consistency globally and locally in two steps. Experiments show that our method outperformed the state-of-the-art unsupervised Re-ID methods.
Ancong Wu, Wei-Shi Zheng 0001, Jian-Huang Lai
ICCV1
2019 Deep Semi-Supervised Person Re-Identification with External Memory
abstract
To overcome the scalability problem of supervised person re-identification (Re-ID), we consider the semi-supervised person Re-ID problem of learning from a limited number of labeled images of a few identities and a large number of unlabeled images. To this end, we propose an external-memory-based deep semi-supervised person Re-ID model (EDS). Based on the external memory, two loss functions are designed so as to effectively cope with the relation between labeled and unlabeled data for overcoming the limitation of batch size in each epoch in deep learning. Therefore, an effective deep semi-supervised learning method can be performed. Extensive experiments validate the superiority of the proposed method for semi-supervised person Re-ID.
Qize Yang, Ancong Wu, Wei-Shi Zheng 0001
ICME2
2019 Deep asymmetric video-based person re-identification
Jingke Meng, Ancong Wu, Wei-Shi Zheng 0001
Pattern Recognit.2
2018 Deep Low-Resolution Person Re-Identification
abstract
Person images captured by public surveillance cameras often have low resolutions (LR) in addition to uncontrolled pose variations, background clutters and occlusions. This gives rise to the resolution mismatch problem when matched against the high resolution (HR) gallery images (typically available in enrolment), which adversely affects the performance of person re-identification (re-id) that aims to associate images of the same person captured at different locations and different time. Most existing re-id methods either ignore this problem or simply upscale LR images. In this work, we address this problem by developing a novel approach called Super-resolution and Identity joiNt learninG (SING) to simultaneously optimise image super-resolution and person re-id matching. This approach is instantiated by designing a hybrid deep Convolutional Neural Network for improving cross-resolution re-id performance. We further introduce an adaptive fusion algorithm for accommodating multi-resolution LR images. Extensive evaluations show the advantages of our method over related state-of-the-art re-id and super-resolution methods on cross-resolution re-id benchmarks.
Jiening Jiao, Wei-Shi Zheng 0001, Ancong Wu, Xiatian Zhu, Shaogang Gong
AAAI3
2018 Adversarial Open-World Person Re-Identification
Xiang Li 0032, Ancong Wu, Wei-Shi Zheng 0001
ECCV (2)2
2018 Letter-Level Writer Identification
abstract
Writer Identification aims to identify a certain writer from a given group of candidates by their handwriting. Although it is very significant in security systems like bank account verification systems, existing works focus on document-level or text-level writer identification. This limits their scalabilities and flexibilities in realistic scenarios as they require complete document or text. To facilitate the realistic applications of writer identification, we propose a novel technology, letter-level writer identification, which requires only a few letters as the identification cue. It is challenging due to large intra-class discrepancy and implicit identifiable writing cues. Considering these challenges, we propose a novel deep model called Multi-Branch Encoding net (Mul-BEnc). To evaluate our model and provide a benchmark for this problem, we have collected a large Letter-stroke sequence Writer identification DataBase (LetWriterDB). The experimental results validate the effectiveness of our model.
Zelin Chen, Hong-Xing Yu, Ancong Wu, Wei-Shi Zheng 0001
FG3
2018 Light Person Re-Identification by Multi-Cue Tiny Net
abstract
Nowadays, person re-identification (re-id) receives much attraction and it is for matching person images across disjoint camera views. Although many methods are developed, none of the state-of-the-art models (especially for the deep models) can be deployed by a camera's chip due to limited memory and weak computation. To address this problem, we propose a framework called Multi-Cue Tiny Net, which is a combination of tiny convolutional neural networks (CNNs) with small model size. And we call the person re-identification based on limited memory and weak computation the Light Person Re-Identification. Three different tiny nets are used to learn the complementary features and four cues of images are used for learning features with different information. All the features will be concatenated to form a fusion feature. Besides, to reduce the dimension of the features, and keep the small model size and high performance accuracy, we pre-trained the tiny nets and then retrained them after adding a fully connected layer to the global average pooling layer. Experimental results on challenging person re-identification dataset show that our approach yields promising accuracy with light memory and computation.
Weiji Wu, Ancong Wu, Wei-Shi Zheng 0001
ICIP2
2018 Does A Body Image Tell Age?
abstract
Age estimation is an important task in computer vision and is widely used in applications. However, such a technology is largely affected by the resolution of face, and it would be a challenge if one has to estimate the age of a person at a distance. While body image of a person is often captured more clearly, when and how to use body-based visual cues for age estimation are largely under studied. In this work, we argue that body-based visual cues are better for estimating the age group and can assist the estimation of exact age value. For this purpose, we develop a Body-based Age Net (BAN) that unifies selective local convolution features and contextual convolution features. The network is designed based on two assumptions: 1) a person's wearing is closely related to his/her age group property; 2) some selective local parts of a body are more discriminative for age group estimation. We have contributed a large-scale and publicly available Body Age (BAG) dataset. We have quantitatively evaluated the proposed model on BAG.
Baoyu Yuan, Ancong Wu, Wei-Shi Zheng 0001
ICPR2
2018 Adversarial Attribute-Image Person Re-identification
abstract
While attributes have been widely used for person re-identification (Re-ID) which aims at matching the same person images across disjoint camera views, they are used either as extra features or for performing multi-task learning to assist the image-image matching task. However, how to find a set of person images according to a given attribute description, which is very practical in many surveillance applications, remains a rarely investigated cross-modality matching problem in person Re-ID. In this work, we present this challenge and leverage adversarial learning to formulate the attribute-image cross-modality person Re-ID model. By imposing a semantic consistency constraint across modalities as a regularization, the adversarial learning enables to generate image-analogous concepts of query attributes for matching the corresponding images at both global level and semantic ID level. We conducted extensive experiments on three attribute datasets and demonstrated that the regularized adversarial modelling is so far the most effective method for the attribute-image cross-modality person Re-ID problem.
Zhou Yin, Wei-Shi Zheng 0001, Ancong Wu, Hong-Xing Yu, Hai Wan, Feiyue Huang, Jian-Huang Lai
IJCAI3
2017 RGB-Infrared Cross-Modality Person Re-identification
abstract
Person re-identification (Re-ID) is an important problem in video surveillance, aiming to match pedestrian images across camera views. Currently, most works focus on RGB-based Re-ID. However, in some applications, RGB images are not suitable, e.g. in a dark environment or at night. Infrared (IR) imaging becomes necessary in many visual systems. To that end, matching RGB images with infrared images is required, which are heterogeneous with very different visual characteristics. For person Re-ID, this is a very challenging cross-modality problem that has not been studied so far. In this work, we address the RGB-IR cross-modality Re-ID problem and contribute a new multiple modality Re-ID dataset named SYSU-MM01, including RGB and IR images of 491 identities from 6 cameras, giving in total 287,628 RGB images and 15,792 IR images. To explore the RGB-IR Re-ID problem, we evaluate existing popular cross-domain models, including three commonly used neural network structures (one-stream, two-stream and asymmetric FC layer) and analyse the relation between them. We further propose deep zero-padding for training one-stream network towards automatically evolving domain-specific nodes in the network for cross-modality matching. Our experiments show that RGB-IR cross-modality matching is very challenging but still feasible using the proposed model with deep zero-padding, giving the best performance. Our dataset is available at http:// isee.sysu.edu.cn/project/RGBIRReID.htm.
Ancong Wu, Wei-Shi Zheng 0001, Hong-Xing Yu, Shaogang Gong, Jian-Huang Lai
ICCV1
2017 Cross-View Asymmetric Metric Learning for Unsupervised Person Re-Identification
abstract
While metric learning is important for Person reidentification (RE-ID), a significant problem in visual surveillance for cross-view pedestrian matching, existing metric models for RE-ID are mostly based on supervised learning that requires quantities of labeled samples in all pairs of camera views for training. However, this limits their scalabilities to realistic applications, in which a large amount of data over multiple disjoint camera views is available but not labelled. To overcome the problem, we propose unsupervised asymmetric metric learning for unsupervised RE-ID. Our model aims to learn an asymmetric metric, i.e., specific projection for each view, based on asymmetric clustering on cross-view person images. Our model finds a shared space where view-specific bias is alleviated and thus better matching performance can be achieved. Extensive experiments have been conducted on a baseline and five large-scale RE-ID datasets to demonstrate the effectiveness of the proposed model. Through the comparison, we show that our model works much more suitable for unsupervised RE-ID compared to classical unsupervised metric learning models. We also compare with existing unsupervised REID methods, and our model outperforms them with notable margins. Specifically, we report the results on large-scale unlabelled RE-ID dataset, which is important but unfortunately less concerned in literatures.
Hong-Xing Yu, Ancong Wu, Wei-Shi Zheng 0001
ICCV2
2017 Correlation Based Identity Filter: An Efficient Framework for Person Search
Wei-Hong Li 0001, Yafang Mao, Ancong Wu, Wei-Shi Zheng 0001
ICIG (1)3
2017 Robust Depth-Based Person Re-Identification
abstract
Person re-identification (re-id) aims to match people across non-overlapping camera views. So far the RGB-based appearance is widely used in most existing works. However, when people appeared in extreme illumination or changed clothes, the RGB appearance-based re-id methods tended to fail. To overcome this problem, we propose to exploit depth information to provide more invariant body shape and skeleton information regardless of illumination and color change. More specifically, we exploit depth voxel covariance descriptor and further propose a locally rotation invariant depth shape descriptor called Eigen-depth feature to describe pedestrian body shape. We prove that the distance between any two covariance matrices on the Riemannian manifold is equivalent to the Euclidean distance between the corresponding Eigen-depth features. Furthermore, we propose a kernelized implicit feature transfer scheme to estimate Eigen-depth feature implicitly from RGB image when depth information is not available. We find that combining the estimated depth features with RGB-based appearance features can sometimes help to better reduce visual ambiguities of appearance features caused by illumination and similar clothes. The effectiveness of our models was validated on publicly available depth pedestrian datasets as compared to related methods for re-id.
Ancong Wu, Wei-Shi Zheng 0001, Jian-Huang Lai
IEEE Trans. Image Process.1
2016 Top-Push Video-Based Person Re-identification
abstract
Most existing person re-identification (re-id) models focus on matching still person images across disjoint camera views. Since only limited information can be exploited from still images, it is hard (if not impossible) to overcome the occlusion, pose and camera-view change, and lighting variation problems. In comparison, video-based re-id methods can utilize extra space-time information, which contains much more rich cues for matching to overcome the mentioned problems. However, we find that when using video-based representation, some inter-class difference can be much more obscure than the one when using still-image-based representation, because different people could not only have similar appearance but also have similar motions and actions which are hard to align. To solve this problem, we propose a top-push distance learning model (TDL), in which we integrate a top-push constrain for matching video features of persons. The top-push constraint enforces the optimization on top-rank matching in re-id, so as to make the matching model more effective towards selecting more discriminative features to distinguish different persons. Our experiments show that the proposed video-based reid framework outperforms the state-of-the-art video-based re-id methods.
Jinjie You, Ancong Wu, Xiang Li 0032, Wei-Shi Zheng 0001
CVPR2
2016 An enhanced deep feature representation for person re-identification
abstract
Feature representation and metric learning are two critical components in person re-identification models. In this paper, we focus on the feature representation and claim that hand-crafted histogram features can be complementary to Convolutional Neural Network (CNN) features. We propose a novel feature extraction model called Feature Fusion Net (FFN) for pedestrian image representation. In FFN, back propagation makes CNN features constrained by the handcrafted features. Utilizing color histogram features (RGB, HSV, YCbCr, Lab and YIQ) and texture features (multi-scale and multi-orientation Gabor features), we get a new deep feature representation that is more discriminative and compact. Experiments on three challenging datasets (VIPeR, CUHK01, PRID450s) validates the effectiveness of our proposal.
Shangxuan Wu, Ying-Cong Chen, Xiang Li 0032, Ancong Wu, Jinjie You, Wei-Shi Zheng 0001
WACV4